Secret AI
A confidential GPU worker network with public VM verification procedures; ordinary SDK calls do not automatically verify attestation.
INPUT / OUTPUT PRIVACY
Stage 1Partial proof
ASSESSED DEPLOYMENT
Direct confidential worker inference
Verification labels describe capabilities established by the cited sources. This review did not perform live cryptographic attestation, reproduce production builds or complete an independent security audit.
Who actually runs the models?
Secret AI workers run the inference engines themselves. WorkerManager registers workers and returns service endpoints; users connect to a worker rather than selecting unrelated upstream API providers. The SDK discovers worker URLs and invokes their locally served models.
Deployment and TEE coverage
Covers direct Secret AI model-worker endpoints on GPU-enabled SecretVM deployments. Applications, fine-tuning, training and the separately described private RAG store require their own assessments. Worker discovery and an API key do not establish a verified inference channel.
The worker architecture combines Intel TDX or AMD confidential VMs with NVIDIA confidential GPUs. SecretVM exposes separate CPU, runtime and GPU evidence. A CPU-only SecretVM or an unchecked worker endpoint is not proof of confidential GPU inference.
Why Stage 1?
Stage requirementsAt least one content layer has independently checkable technical evidence, including partial proof within a layer. The full privacy boundary is incomplete. The stage tooltip distinguishes partial findings from fully met requirements for Stage 2.
STEPS FOR STAGE 2
1 of 3 requirements for Stage 2 met. Missing proof: Request path, Logs & storage.
- Request pathNot established
Private inputs and outputs across the request path, with E2EE or an equivalent attested channel that protects against intermediaries and binds keys to the accepted workload.
Documented only. Attestation is not automatic
- Inference executionMet
Protected inference with a verifiable workload identity and protection covering the CPU, GPU and every place content is processed.
Technically verifiable. CPU + GPU; manual checks
- Logs & storageNot established
No content logging, persistent storage, human review or training. Any temporary content-derived cache must have verifiable isolation and enforced removal.
Unknown. Cache removal not established
Identity and operational metadata are assessed separately below. They do not set the content stage; prompt or response content in telemetry remains part of the content-retention assessment.
Request path
Documented onlySDK verification gap. SecretVM binds its attestation web server's TLS certificate fingerprint into report_data. The manual verification guide checks this binding. However, the reviewed ChatSecret.get_attestation() method returns an empty object and is marked unimplemented; regular SDK inference does not establish automatic attestation verification or pinning of the actual inference-port certificate. Verify that endpoint and its key separately before sending content.
Inference execution
Technically verifiableWorker evidence available. Published procedures let a reviewer reproduce the VM's expected measurements from firmware, kernel, root filesystem and Docker Compose configuration. GPU workers also expose NVIDIA evidence for separate validation. This earns partial proof for an inspectable worker deployment; it is not evidence that every discovered worker or serving GPU has passed the user's checks.
Logs & storage
UnknownWorker-dependent. The reviewed architecture and SDK do not establish enforced content-log suppression, bounded KV/prefix-cache removal, backup behavior or a no-training guarantee for every worker. An attested VM can still run application code that stores content. The separate PRAG design intentionally stores confidential retrieval data and must not be treated as stateless inference.
Identity & metadata
Documented onlyToken usage reported. Workers validate subscriptions and report input/output token counts to RewardsManager for payment distribution. These operational records are distinct from content privacy. The reviewed sources do not establish a complete retention schedule or unlinkability across subscriptions, worker connections and payments.
Remaining trust assumptions
- Intel or AMD VM isolation and NVIDIA GPU isolation, with both evidence chains checked for the selected worker.
- WorkerManager discovery, the chosen worker's measured application configuration and its update policy.
- The user's verifier and binding of the inference connection to the approved worker, beyond ordinary SDK authentication.
Limits of this assessment
- The SDK's attestation accessor is a stub in the reviewed source; an API key is not attestation.
- VM measurement verification alone does not establish GPU state, model identity or content deletion.
- No live worker quote, serving connection or independent deployment audit was verified in this review.
Primary-source record
5 sources · reviewed 2026-10-06
Sources reviewed 2026-10-06