Home

/

Blog Details

On-prem AI and sovereign AI: anatomy of a customer deployment

On-prem AI, sovereign AI: servers, GPUs, vLLM, model, RAG, gateway and observability. The six layers of a customer-site deployment, and their pitfalls.

October 1, 2026

We announced our partnership with Dell on Tuesday. Since then, I've been asked the same question three times: "So what does an AI running at the customer's site actually look like?"

‍

It's a good question. People talk a lot about "sovereign AI" and "on-prem AI", but you rarely see what's inside. So here is the full anatomy, layer by layer, from the hardware to the use case.

‍

1. The foundation: servers and GPUs

‍

It all starts with hardware: a Dell server and NVIDIA GPUs.

‍

The only real trap at this level is sizing. Start from the actual workload, not the available budget. Three questions matter:

‍

  • How many concurrent users?
  • How long are the contexts? (long documents, conversation history)
  • What response time is acceptable?

‍

An undersized GPU doesn't make a slow project. It makes an abandoned one. Users don't come back to a tool that keeps them waiting.

‍

2. The inference engine: vLLM

‍

A graphics card on its own isn't enough. You need a layer that turns it into a service, and that's the job of the inference engine. We use vLLM, which handles request batching, memory and throughput.

‍

This is the layer where a demo for one user becomes a service used by two hundred people at the same time. Many projects that "worked fine in the POC" fail right here.

‍

3. The model: Qwen by default

‍

Our default model is Qwen. It has open weights, it's hosted on the customer's servers, and it makes no outbound calls. The data never leaves the infrastructure.

‍

We pick the model for the task, not for this month's leaderboard. Benchmarks change every week, while a production use case needs to stay stable.

‍

One rule to remember: an average model served well beats an excellent model served badly.

‍

4. The data: RAGFlow and PostgreSQL

‍

This is the least glamorous layer, and also the one that matters most.

‍

  • RAGFlow handles document ingestion and search.
  • PostgreSQL stores structured and application data, with a semantic layer on top.

‍

The quality of a document assistant doesn't depend on the choice of LLM. It comes down to two less visible things: how documents are chunked and how access rights are managed. Poor chunking gives off-target answers. Poorly managed rights give you an assistant you can't open up to the whole company.

‍

5. The front door: a gateway (LiteLLM)

‍

Every request goes through a gateway. We use LiteLLM, which provides:

‍

  • routing between models;
  • quotas;
  • traceability;
  • chargeback by department.

‍

Without a gateway, nobody knows who's using what. After six months, the project can no longer be managed: there's no way to justify costs, arbitrate between teams or plan for growth.

‍

6. Guardrails and observability: NeMo Guardrails and Langfuse

‍

Two components complete the architecture:

‍

  • NeMo Guardrails to filter inputs and outputs;
  • Langfuse for observability: traces, latency, answer quality.

‍

The logic is simple. Without traces, you can't diagnose. Without diagnosis, you can't improve.

‍

At the end of the chain: a use case

‍

All these layers exist only to run a use case, for example:

‍

  • document search across a technical knowledge base;
  • support for customer service teams;
  • data extraction from incoming documents.

‍

Nothing spectacular, but these are tools that run every day and that teams actually use.

‍

What the diagram doesn't show

‍

An architecture fits in a diagram. What takes the most time is what you don't see in it:

‍

  • taking over the existing document base: mixed formats, duplicates, outdated versions;
  • integration with the information system: directory, access rights, business applications;
  • operations once in production: updates, monitoring, changing usage.

‍

That's exactly the work we do with the integrators who deploy these architectures for SMEs and mid-sized companies.

‍

Are you an integrator looking to offer on-site AI to your customers? Let's talk.

Ready to accelerate
your business with AI?

Get in touch and find a solution.

Contact Us

Contact Us

Collaborate with our team

Organize AI projects at a glance