RESEARCHED PROVIDER · PRIMARY SOURCES

Tinfoil

Hosted open-weight model inference with client verification of confidential workloads and their encrypted channels.

INPUT / OUTPUT PRIVACY

Stage 1

Partial proof

ASSESSED DEPLOYMENT

Inference API with verifying SDK

RESEARCH REVIEW

2026-10-06 · primary-source review

Provider website

Verification labels describe capabilities established by the cited sources. This review did not perform live cryptographic attestation, reproduce production builds or complete an independent security audit.

Who actually runs the models?

Tinfoil operates model-loaded enclaves with a vLLM inference server and a confidential model router. Using rented GPU infrastructure does not make the service an upstream model-API aggregator.

Deployment and TEE coverage

Assesses inference-only API calls through Tinfoil's verification path with a current SDK and client-held cache secret. Raw HTTP without the secret has weaker cache isolation. Chat backups, customer Containers and optional web-search recipients have separate data lifecycles.

The documented API routes through attested enclaves. GPU workers verify NVIDIA confidential mode during boot; model-weight volumes are checked against integrity commitments. This rating requires a verifying client and excludes external tools.

At least one content layer has independently checkable technical evidence, including partial proof within a layer. The full privacy boundary is incomplete. The stage tooltip distinguishes partial findings from fully met requirements for Stage 2.

STEPS FOR STAGE 2

2 of 3 requirements for Stage 2 met. Missing proof: Logs & storage.

  1. Request pathMet

    Private inputs and outputs across the request path, with E2EE or an equivalent attested channel that protects against intermediaries and binds keys to the accepted workload.

    Independently auditable. Pinned code and channel keys

  2. Inference executionMet

    Protected inference with a verifiable workload identity and protection covering the CPU, GPU and every place content is processed.

    Independently auditable. Code + model-weight integrity

  3. Logs & storageNot established

    No content logging, persistent storage, human review or training. Any temporary content-derived cache must have verifiable isolation and enforced removal.

    Documented only. Protected cache; no expiry

Identity and operational metadata are assessed separately below. They do not set the content stage; prompt or response content in telemetry remains part of the content-retention assessment.

Request path

Independently auditable

Attested SDK path. SDK verification binds the connection to an enclave-generated key before content is sent. Public code, reproducible images and authenticated release pins connect the accepted workload to its source. The confidential router verifies downstream model enclaves; the client need not trust a plaintext host gateway.

Inference execution

Independently auditable

Reproducible workload. The backend documents measured configuration, reproducible CVM images and dm-verity model-weight commitments. The measured boot code verifies GPU confidential mode and aborts on failure. GPU verification is transitive through the CPU workload, not an independently fetched per-request GPU report.

Logs & storage

Documented only

Cache removal gap. Read-only disks prevent disk persistence, but do not establish no post-request content-derived state. Tinfoil documents isolated GPU prefix caches with client-held secrets, no explicit expiry, and eviction under resource pressure or enclave shutdown. Its no-content-retention policy does not establish a removal deadline or independently checkable per-entry cleanup. Isolation is evidenced; an enforced cache-removal lifecycle sufficient for Stage 2 was not established.

Identity & metadata

Documented only

Operational records. Policy retains account, billing, usage and security data. Account deletion can take 30 days across backups; billing records are typically kept up to seven years. IP/security logs are handled through Clerk, without a fixed duration here. The schedule and its exceptions are policy evidence, not independently verified deletion enforcement.

Remaining trust assumptions

  • Hardware and firmware roots, including the GPU's confidential-computing verifier.
  • Reviewed router and inference code, the SDK, source-build provenance and accepted release pins.
  • External tools have their own recipients and policies; they are outside inference-only confidentiality.

Limits of this assessment

  • Open source and attestation enable auditing; they do not mean this index completed an independent security audit.
  • Content confidentiality does not provide anonymous access or erase billing metadata.
  • Stateless disks do not mean no prompt cache; cache entries have no explicit expiry.
  • API content retention and encrypted Chat backup retention are different guarantees.

Primary-source record

6 sources · reviewed 2026-10-06

Back to providers

Sources reviewed 2026-10-06