RESEARCHED PROVIDER · PRIMARY SOURCES

Privasys

Hosted and dedicated confidential inference fleets using open deployment software and attestation-bound TLS.

INPUT / OUTPUT PRIVACY

Stage 1

Partial proof

ASSESSED DEPLOYMENT

Verified RA-TLS inference API

RESEARCH REVIEW

2026-10-06 · primary-source review

Provider website

Verification labels describe capabilities established by the cited sources. This review did not perform live cryptographic attestation, reproduce production builds or complete an independent security audit.

Who actually runs the models?

Privasys runs open-weight inference fleets and supplies confidential-ai and Enclave OS software for dedicated deployments. vLLM and the model weights execute locally inside a GPU-enabled confidential VM; the assessed service is not a catalog routing requests to unrelated model APIs.

Deployment and TEE coverage

Covers a fleet's direct confidential-ai API with a client that verifies and pins its RA-TLS evidence. Privasys Chat, Drive, stored conversations, RAG, agents and connectors extend the boundary and are excluded. A dashboard that fetches evidence on a user's behalf is not by itself verification of the user's own inference connection.

The documented runtime uses Intel TDX and NVIDIA confidential-compute GPUs. Its API exposes model and workload measurements. Production image policy, GPU evidence, trusted model-disk hashes and each fleet's actual configuration must be checked; launching the standalone proxy outside Enclave OS does not create a TEE.

At least one content layer has independently checkable technical evidence, including partial proof within a layer. The full privacy boundary is incomplete. The stage tooltip distinguishes partial findings from fully met requirements for Stage 2.

STEPS FOR STAGE 2

2 of 3 requirements for Stage 2 met. Missing proof: Logs & storage.

  1. Request pathMet

    Private inputs and outputs across the request path, with E2EE or an equivalent attested channel that protects against intermediaries and binds keys to the accepted workload.

    Technically verifiable. Direct verifying clients only

  2. Inference executionMet

    Protected inference with a verifiable workload identity and protection covering the CPU, GPU and every place content is processed.

    Technically verifiable. Runtime + model-disk checks

  3. Logs & storageNot established

    No content logging, persistent storage, human review or training. Any temporary content-derived cache must have verifiable isolation and enforced removal.

    Documented only. No bounded removal established

Identity and operational metadata are assessed separately below. They do not set the content stage; prompt or response content in telemetry remains part of the content-retention assessment.

Request path

Technically verifiable

Pinned RA-TLS. The RA-TLS certificate carries a TDX quote binding its public key and a challenge nonce into report_data. Clients can compare runtime measurements, configuration hashes and model-disk identity before sending API content. This rating requires a fresh challenge and verification of the actual connection, not ordinary HTTPS or a server-side dashboard result.

Inference execution

Technically verifiable

Measured local inference. The architecture places the proxy, manager and local vLLM engine inside a TDX guest with confidential GPU memory. Attestation exposes boot/runtime measurements and a model-disk digest; the public runtime source supplies deployment inputs. The service describes GPU evidence in its trust chain. A reviewer still needs to verify the selected fleet's complete CPU/GPU chain and approved code/configuration; output replay alone does not prove privacy.

Logs & storage

Documented only

Cache lifecycle gap. The privacy documentation states that prompts and responses are not logged or written to cleartext disk and that KV state stays in protected memory. It does not establish enforced cache expiry or clearing for the assessed fleet. Its reproducible-image guarantee depends on the operator publishing all build inputs. Chat/RAG separately persist conversations, embeddings and feedback; those products are outside this inference-only scope.

Identity & metadata

Documented only

Traffic and resource counters. The host can observe connection size and duration. The manager exports resource counters and model identity to the management plane. The reviewed privacy documentation describes these observations but does not establish a complete account, IP or billing-retention schedule for each fleet.

Remaining trust assumptions

  • Intel TDX and NVIDIA confidentiality and attestation roots, plus the client's trusted collateral providers.
  • The fleet's published build inputs, approved image/configuration hashes and trusted model-disk signature.
  • The application's verifier and update policy; arbitrary tools or connectors can expand the privacy boundary.

Limits of this assessment

  • An ordinary OpenAI-compatible HTTP client does not automatically verify RA-TLS evidence.
  • Reproducible release inputs and deletion behavior must be established for the chosen fleet.
  • This rating covers direct inference, not stored chat/RAG data or arbitrary agents; no live fleet was cryptographically verified.

Primary-source record

6 sources · reviewed 2026-10-06

Back to providers

Sources reviewed 2026-10-06