Sovereign artificial intelligence
Language models that run on your premises or on Swiss infrastructure, with sourced answers and verifiable governance. For sectors where sending a document to an American API is not an option.
Why self-host
Three reasons come up again and again with our clients. Compliance first: the Swiss FADP governs the disclosure of personal data abroad, and American AI APIs make that assessment considerably harder for healthcare, fiduciary work, pharma and the public sector.
Then trade secrets. A contract, a patent or a client file has no business in a third-party service's prompt, whatever its published terms say.
And finally control. A self-hosted model does not change behaviour overnight because a vendor updated its API. Your evaluations stay valid and your costs stay predictable.
What we build
We do not sell 'AI'. We address a named use case with a measurable success criterion agreed before we start.
Document search
Query thousands of internal documents in natural language, with the source and originating page cited every time.
Structured extraction
Turn invoices, contracts or reports into usable data, with an error rate that is measured rather than assumed.
Business agents
Automate a chain of tasks across your existing systems, with human approval at the sensitive steps.
Classification and triage
Route enquiries, correspondence or tickets automatically, according to your own categories.
RAG before fine-tuning
The reflex of 'we need to train the model on our data' is almost always premature. Retrieval-augmented generation, or RAG, brings business knowledge in at query time. It stays current without retraining and, above all, allows sources to be cited.
In a regulated context, an answer with no verifiable reference is unusable. So we reserve fine-tuning for cases where the output format or the tone must be deeply adapted, and never for injecting knowledge.
Evaluate before deploying
This is where AI projects most often fail. Without a business test set, there is no way to know whether the system regresses when the model, the prompt or the document base changes.
So we build an evaluation set with your teams before writing the first line of integration: real questions, expected answers, quantified acceptance criteria. That document is what makes it possible to declare the system ready on something other than intuition.
Compliance and validated environments
An internal LLM is an infrastructure component like any other: isolated network or hosting on your premises, role-based access control, request logging, and registration of the processing activity in your FADP record.
For validated environments, in pharma particularly, qualification sits within the GAMP 5 and Annex 11 framework, with a reproducible test file. We also prepare the documentation the EU AI Act expects for systems classified as limited or high risk.
The hardware: less than you would think
The right approach is not to aim for the largest model but for the smallest one that succeeds at your task. A model of 8 to 30 billion parameters, quantised where appropriate, runs on a single GPU server and is enough for the large majority of enterprise use cases.
We size from your evaluation set and your real load, not from a public leaderboard. Where the hardware investment is not justified, we deploy on Swiss infrastructure at Infomaniak or Exoscale, which settles the data residency question without tying up capital.
An example
Company-management platform with an AI copilot: shareholder registry, general meetings and intelligent search across cantonal tax sources, with an absolute requirement of zero hallucination on Swiss tax law. RAG architecture, LangGraph, sovereign hosting.
Frequently asked questions
What hardware does an internally hosted LLM need?
Most often a single server with a professional GPU is enough for a model of 8 to 30 billion parameters serving a few dozen users. We size from your real load and your latency requirements, after measuring on your evaluation set.
RAG or fine-tuning?
RAG in more than nine cases out of ten. It brings knowledge in at query time, stays current without retraining and allows sources to be cited. Fine-tuning is only justified to adapt an output format or a very particular tone.
Is hosting in Switzerland enough for the FADP?
It is an important element but not sufficient on its own. You also need to document the processing in your register and define access rights, the retention period for queries and how data subjects are informed. We deliver that documentation with the system.
How do you prevent hallucinations?
By requiring source citation, restricting the model to the supplied corpus, and measuring the error rate against a business test set. Where a wrong answer is expensive, the system must be able to answer that it does not know. That is an acceptance criterion, not a detail.
How long until a first system in production?
Allow six to twelve weeks for a document assistant on a defined corpus, evaluation set included. Assembling and cleaning the corpus often takes longer than the technical integration.
Let's talk about your project.
Describe your need in a few lines. We answer within 48 working hours, with a first technical opinion rather than a sales pitch.