Solution
Private AI infrastructure
An assistant or an agent is only as good as the infrastructure serving it: the available memory decides the model, the context decides which documents it reads, the gateway decides what leaves. I size and deploy the infrastructure that serves the model, on a server of yours, with a Swiss host or behind a gateway to a cloud provider's APIs, and I measure latency, cost and quality on your own cases.
A model served, measured and governed
Sizing comes before buying
Memory, context, throughput and concurrent users are measured on a candidate model before a machine is ordered or a subscription signed. A ten-user problem doesn’t justify a five-card workstation.
Cost per request is known
On-site model, Swiss hosting and cloud API are compared on the same documents: latency, quality on an approved test set, cost per request, hardware, energy and operations included.
Every outgoing call is logged and capped
A gateway applies the rules per data type, logs every call and caps spend per team and per day. The cloud provider receives what the rule allows, nothing else.
The model is swapped without rewriting the application
Assistants and agents talk to a stable interface. Changing model, or moving from the cloud to your machines, becomes an operations task with its tests, not a new project.
In practice
- Sizing: candidate model, memory, context length, throughput and concurrent users measured on your documents before any purchase
- An inference server on your machines or with a Swiss host, vLLM or an equivalent service for instance, with model updates tested like a production release
- A gateway to a cloud provider's AI APIs: call logging, spend caps, rules per data type and a stable interface for the applications
- Vector search and document ingestion: chunking, metadata, permissions and scheduled reindexing
- An approved evaluation set and regular measures of quality, latency and cost, with the results recorded
- Documentation of location, access, retention and training use, service by service
Systems involved
- GPU servers, dedicated workstations or Swiss hosting
- Open models and cloud providers' APIs, behind one gateway
- PostgreSQL with vector search and the document stores in place
- The assistants, agents and extraction pipelines that consume the model
- Existing monitoring, logs and backups
Service lineArtificial intelligence →
The documents the model has to read
The same work, against each sector’s own constraints. Every card opens the full sector.
Healthcare and life sciences
A model served on your premises for the documents that never leave them
The open model runs on a machine of the clinic or the laboratory, sized to your volumes, with latency, cost and quality measured on your own cases before the assistant reads a record.
Manufacturing
An assistant over the technical documentation, served inside the plant
The open model runs on a server in the plant and reads the manuals, the procedures and the maintenance history without a document leaving the site. Sizing comes before the machine is bought.
How it runs
Measurement
Two or three candidate models run on your documents and your test set, on a trial machine or an API, with quality, latency, memory and cost recorded.
Decision
Management chooses between your machines, a Swiss host and a cloud API on the measured figures and on the documented contract, location and retention.
Deployment
The inference server or the gateway is deployed, described in code, with logs, caps and per-data-type rules in place before the first user.
Operations
Model updates go through the test set, the measures continue and the figures enter the monthly report.
To go deeper

Kimi K3, Soofi S and Apertus: an open-model snapshot
A dated snapshot, not a permanent ranking: release status, weight precision, internal benchmarks, total cost and deployment conditions.

Local AI: three machine profiles for several kinds of work
Mac mini, DGX Spark or a Blackwell workstation: three starting points for on-site AI, with the memory, context and operational assumptions that still need testing.

Internal AI without a US cloud: realistic in 2026?
Open models, Swiss hosting and local deployment: a dated method for selecting internal AI without reducing sovereignty to the server's location.
Test the fit: Private AI infrastructure
Describe the context, constraints and decision you need to make. The first conversation qualifies scope, boundaries and the next useful step.
Describe the situation