Privacy & Data
Varen's data-sovereignty model is about where your content is stored. Your durable content — conversations, uploaded files, knowledge base, profile, and indexes — is stored on your node, not in Varen's database. This reference explains what stays on your node, what the control plane stores, and how content is handled in transit when you ask a question.
The architecture
Your durable content lives on your node. The Varen control plane (the cloud component) brokers connections, handles authentication, routes work, coordinates model access, and manages billing. It is not the durable store for your conversations, files, or knowledge base — that data stays on the hardware you've designated as your node.
This separation is architectural, not a policy setting. Your content isn't kept out of a Varen database by a configuration toggle or a promise about behaviour — the durable store simply is your node. The control plane holds only the operational records it needs to run the service.
What the control plane stores
The Varen control plane persists a defined, limited set of information:
- Your account details — email address, plan, and credit balance
- Node connection state — whether a node is online and which account it belongs to
- Billing records — usage amounts and timestamps, not the content of conversations
It does not store your conversations, uploaded files, knowledge base, or session history as durable customer content. Those remain on your node.
Content in transit during a query
Answering a question is different from storing data. To produce a reply, the control plane assembles the prompt and calls the model provider, so it necessarily handles some of your content in transit for that turn:
- The text of your question, as sent for the current turn
- The relevant context retrieved from your knowledge base for that turn
- The bytes of an uploaded file when it is extracted or analysed — SaaS fetches the file from the node and forwards it to the extraction or model processor (see File extraction and the tools rail)
This is expected and necessary for inference: it is how the answer gets generated. The key point is that SaaS passes this content through to fulfil the request rather than retaining it as your durable store. Where a hosted model is used — or where an extracted page is escalated to a hosted vision provider — that content also reaches the provider, whose retention behaviour is governed by its own terms.
The encrypted tunnel
Your node connects to the Varen control plane through an encrypted tunnel. Traffic between your browser and your node travels through this tunnel — your browser connects to varen.tech, and varen.tech relays the connection through to your node. WireGuard, gRPC, and TLS protect this traffic in transit. As noted above, this transport encryption is separate from inference: on the inference and extraction paths SaaS terminates the request and handles the content to build and route the prompt.
This architecture means your node doesn't need to be directly reachable from the internet. The encrypted tunnel provides the connection path. Your node reaches out to the control plane; the control plane holds the connection open; your browser accesses the node through that connection.
Hosted inference providers
When you use a hosted model, the in-transit content described above doesn't stop at Varen — it continues to the AI provider that generates the reply. The text of your question, along with the relevant retrieved context, is sent to that provider as part of each request.
Varen sends each request independently rather than using the provider as your conversation store, but provider-side logging, retention, and training behaviour are governed by that provider's own terms. This is the same pattern as any cloud-based AI service. Varen does not currently offer fully local answer generation: every reply is produced by a hosted model. What runs locally on your node is the supporting machinery — embeddings, retrieval, indexing — not the final inference. The Technical Overview covers this boundary in detail.
If you have concerns about specific sensitive content when using a hosted provider, note that file content is included in a request when it's directly relevant to the question, and choose accordingly.
File extraction and the tools rail
Extractable files — PDFs, images, and office documents — go through an extraction step that turns their bytes into readable, indexable text. Where that conversion happens depends on the page, and it defines three tiers of data sovereignty:
- Your node. Text-native files (plain text, Markdown, HTML, structured documents) are read directly on your node, and embeddings and indexes are always built and stored there. This content never leaves your hardware for extraction.
- Varen-operated tools. PDF pages that carry a usable embedded text layer are converted by Varen's own tools service — mechanical text parsing, not AI: no models, no OCR, no third party. The document reaches the tools service through the control plane, in transit, for the conversion call only. The tools service is stateless: it does not store request bodies, keep files, or run job queues — pages are processed in memory and discarded when the call completes. Processed text returns to your node, where the fragments and indexes live durably.
- External providers. Pages with no usable text layer — scanned image pages and vector artwork such as charts — are the only extraction content that can reach a hosted vision provider, and never automatically. They wait on the asset page until you choose, page by page, which of them to send. Only your selected pages are rendered and sent; the provider's retention behaviour is governed by its own terms, as with hosted inference.
Choosing no pages is a valid final state: the document stays searchable from its text pages and nothing further is spent or sent. Blank pages are skipped honestly and reach neither the tools rail nor any provider. Archiving an asset cancels any pages still waiting, with no provider calls.
Outcome ledger derivation
Outcome workspaces can turn an application form you upload into a guided interview ledger. You drop the form into [The Outcomes] zone; the form itself is stored on your node, hidden from ordinary indexing, and never enters your knowledge base. Derivation is the one process that reads it, and it crosses two boundaries:
- Varen-operated tools. The form's bytes transit the control plane to Varen's tools service for mechanical structure extraction — the same stateless rail as file extraction: no models, no storage, processed in memory and discarded when the call completes.
- Hosted inference provider. The extracted structure (questions, sections, field types — not the raw file) is converted into a ledger by a hosted model through the same LiteLLM path as other inference. Provider retention behaviour is governed by its own terms, as with hosted inference.
The derived ledger returns to your node and is stored there only — a hidden file beside the source form. The control plane never stores it. Nothing runs against the ledger until you review it in the workspace and explicitly confirm it; a derived draft you do not confirm is never used. Derivation is metered as a single flat fee (credit-gated like inference), and its internal model call is never charged separately.
Node hosting and privacy
Where you run the node determines who controls the hardware your durable content is stored on. The storage-sovereignty model is the same in each case — your durable content lives on the node — but physical and operational control differ.
Your hardware — the node runs on hardware you own, in your home or office. Your durable content is under your physical and operational control, and Varen has no physical or network access to the machine. This is the highest level of hardware sovereignty.
Your cloud account — the node runs on a cloud host provisioned under your own account (AWS, Hetzner, DigitalOcean, or similar). The infrastructure is yours; Varen has no access to it. Your cloud provider has the usual access that comes with operating the underlying infrastructure — the same access any cloud provider has to any tenant's data.
Managed hosting — if you'd rather not run the infrastructure yourself, managed hosting may be available on request. Because topology, operator access, and backup arrangements matter for sovereignty, confirm those specifics with us before relying on this option rather than assuming a particular isolation guarantee.
If regulatory or professional requirements mean you need complete hardware sovereignty — your equipment, your data centre, no exceptions — the self-hosted path is the right one. Contact hello@varen.au if you need guidance on the self-hosted deployment.