Tinfoil
Hosted open-weight model inference with client verification of confidential workloads and their encrypted channels.
INPUT / OUTPUT PRIVACY
Stage 1Partial proof
ASSESSED DEPLOYMENT
Inference API with verifying SDK
Verification labels describe capabilities established by the cited sources. This review did not perform live cryptographic attestation, reproduce production builds or complete an independent security audit.
Who actually runs the models?
Tinfoil operates model-loaded enclaves with a vLLM inference server and a confidential model router. Using rented GPU infrastructure does not make the service an upstream model-API aggregator.
Deployment and TEE coverage
Assesses inference-only API calls through Tinfoil's verification path with a current SDK and client-held cache secret. Raw HTTP without the secret has weaker cache isolation. Chat backups, customer Containers and optional web-search recipients have separate data lifecycles.
The documented API routes through attested enclaves. GPU workers verify NVIDIA confidential mode during boot; model-weight volumes are checked against integrity commitments. This rating requires a verifying client and excludes external tools.
Why Stage 1?
Stage requirementsAt least one content layer has independently checkable technical evidence, including partial proof within a layer. The full privacy boundary is incomplete. The stage tooltip distinguishes partial findings from fully met requirements for Stage 2.
STEPS FOR STAGE 2
2 of 3 requirements for Stage 2 met. Missing proof: Logs & storage.
- Request pathMet
Private inputs and outputs across the request path, with E2EE or an equivalent attested channel that protects against intermediaries and binds keys to the accepted workload.
Independently auditable. Pinned code and channel keys
- Inference executionMet
Protected inference with a verifiable workload identity and protection covering the CPU, GPU and every place content is processed.
Independently auditable. Code + model-weight integrity
- Logs & storageNot established
No content logging, persistent storage, human review or training. Any temporary content-derived cache must have verifiable isolation and enforced removal.
Documented only. Protected cache; no expiry
Identity and operational metadata are assessed separately below. They do not set the content stage; prompt or response content in telemetry remains part of the content-retention assessment.
Request path
Independently auditableAttested SDK path. SDK verification binds the connection to an enclave-generated key before content is sent. Public code, reproducible images and authenticated release pins connect the accepted workload to its source. The confidential router verifies downstream model enclaves; the client need not trust a plaintext host gateway.
Inference execution
Independently auditableReproducible workload. The backend documents measured configuration, reproducible CVM images and dm-verity model-weight commitments. The measured boot code verifies GPU confidential mode and aborts on failure. GPU verification is transitive through the CPU workload, not an independently fetched per-request GPU report.
Logs & storage
Documented onlyCache removal gap. Read-only disks prevent disk persistence, but do not establish no post-request content-derived state. Tinfoil documents isolated GPU prefix caches with client-held secrets, no explicit expiry, and eviction under resource pressure or enclave shutdown. Its no-content-retention policy does not establish a removal deadline or independently checkable per-entry cleanup. Isolation is evidenced; an enforced cache-removal lifecycle sufficient for Stage 2 was not established.
Identity & metadata
Documented onlyOperational records. Policy retains account, billing, usage and security data. Account deletion can take 30 days across backups; billing records are typically kept up to seven years. IP/security logs are handled through Clerk, without a fixed duration here. The schedule and its exceptions are policy evidence, not independently verified deletion enforcement.
Remaining trust assumptions
- Hardware and firmware roots, including the GPU's confidential-computing verifier.
- Reviewed router and inference code, the SDK, source-build provenance and accepted release pins.
- External tools have their own recipients and policies; they are outside inference-only confidentiality.
Limits of this assessment
- Open source and attestation enable auditing; they do not mean this index completed an independent security audit.
- Content confidentiality does not provide anonymous access or erase billing metadata.
- Stateless disks do not mean no prompt cache; cache entries have no explicit expiry.
- API content retention and encrypted Chat backup retention are different guarantees.
Primary-source record
6 sources · reviewed 2026-10-06
Sources reviewed 2026-10-06