Souveräne KI on-premise
Sovereign AI on-premise
„On-premise“ ist gerade ein Verkaufswort. Ich verkaufe es nicht — ich betreibe es. In meinem Keller in Ostfriesland steht ein Server, auf dem Sprachmodelle, Embeddings und Bildgenerierung laufen, ohne dass ein Byte das Haus verlässt. Diese Seite sagt, wann sich das für euch lohnt — und ehrlich, wann nicht.
“On-premise” is a buzzword right now. I don't sell it — I run it. In my basement in East Frisia sits a server running language models, embeddings and image generation without a single byte leaving the building. This page says when that's worth it for you — and honestly, when it isn't.
Wann on-premise sinnvoll ist
- Die Daten dürfen nicht raus. Verträge mit Automotive-, Marine- oder Defence-Kunden, die Cloud ausschließen. Rohlogs, Konstruktionsdaten, Prüfprotokolle. Wenn „schick es an einen US-Anbieter“ keine Option ist, lautet die Frage nicht ob lokal, sondern wie.
- Kontinuierliche Grundlast. Wer dauerhaft inferiert, zahlt in der Cloud jeden Token. Eigene Hardware ist nach der Anschaffung fast nur noch Strom.
- Latenz und Offline. Produktionsnah, ohne Internet-Abhängigkeit, ohne „Ihr Guthaben ist aufgebraucht“.
Wann nicht — die Grenzen
Das ist das eigentliche Verkaufsargument, weil es kaum jemand ausspricht:
- Nur gelegentliche Nutzung? Dann ist eine API billiger als jeder Server. Ein Kasten, der zu 95 % Leerlauf hat, rechnet sich nicht.
- Ihr braucht immer das absolut stärkste Frontier-Modell? Das läuft lokal nicht. Dann ist der richtige Weg ein Frontier-Anbieter (z. B. Anthropic Claude) — wenn nötig mit EU-Region und Auftragsverarbeitungsvertrag.
- Stark schwankende Last (mal nichts, mal Riesenpeak)? Cloud skaliert das besser als gekaufte Hardware, die die meiste Zeit döst.
- Niemand kann den Betrieb übernehmen? Ohne jemanden, der Treiber, Updates und Überwachung macht, wird on-prem zur Dauerbaustelle. Genau da komme ich rein.
On-premise ist kein Glaubensbekenntnis. Oft ist die richtige Antwort ein Hybrid: das lokale Modell für die sensiblen Rohdaten, ein Frontier-Modell für das schwere Reasoning — genau die Architektur, die bei CARIAD produktiv läuft.
Was es real kostet — am Beispiel Zuse
Kein Prospekt, mein eigener Server: eine gebrauchte HP Z840 plus drei RTX 3090 in einem externen Rig — zusammen 72 GB GPU-Speicher mit voller Bandbreite. Zwei Karten fahren das Sprachmodell (Qwen3.6-27B, 256K Kontext, ~65 Token/s), die dritte die Bildgenerierung. Mehr dazu im HandsOn zu Zuse.
Anschaffung, Gebrauchtmarkt — die echten Zahlen von Zuse:
| Posten | grob |
|---|---|
| 3× RTX 3090 (gebraucht, à ~900 €) | ~2.700 € |
| HP Z840 Workstation (gebraucht, mit 128 GB RAM) | ~680 € |
| Externes Rig, Riser-Kabel, zweite GPU-PSU | ~330 € |
| Summe für 72 GB VRAM @ 936 GB/s | ~3.710 € |
Kleiner anfangen geht: gebrauchte Workstation plus eine 3090 und NVMe liegt bei ~1.900–2.100 € — genug für Modelle bis ~14B komfortabel, das 27B mit gekürztem Kontext. Einordnung 2026: Durch die aktuelle Speicherkrise sind RAM, SSD und GPUs teuer; eine gebrauchte 3090 liegt inzwischen bei ~1.000–1.200 €.
Strom (Annahme 0,30 €/kWh): Eine 3090 zieht bis ~350 W, per Power-Cap gedeckelt auf ~250 W. Leerlauf rund um die Uhr ~190 W, drei Karten unter Voll-Last (gedeckelt) ~900 W. Real, je nach Auslastung, grob 500–1.100 € im Jahr (ca. 40–95 € im Monat).
Betrieb ist der Posten, den Prospekte verschweigen: Treiber, Modell-Updates, Überwachung, die Lektionen, die wehtun (Modelle gehören auf SSD, Whisper halluziniert auf Stille). Das kostet keine Euro, sondern Zeit — und ist der Grund, warum die meisten den Betrieb abgeben.
Break-even: Für eine dauerhafte Grundlast unterbietet ein solcher Server jede vergleichbare Dauer-Miete in der Cloud deutlich. Für ein paar Anfragen pro Woche niemals. Zwischen diesen beiden Polen verläuft die Entscheidung — und die rechne ich mit euch konkret durch, nicht mit einer Faustregel aus dem Prospekt.
Modellauswahl
Was lokal läuft, hängt an der Hardware. Auf zwei RTX 3090 (48 GB VRAM) läuft ein 27B-Modell in int4 mit langem Kontext flüssig (~65 tok/s) — genug für RAG, Extraktion, Agenten-Schritte und Klassifikation. Für das schwerste Reasoning greift man zum Frontier-Modell. Die eigentliche Arbeit ist die Aufteilung: was muss lokal bleiben, was darf — anonymisiert — nach draußen.
DSGVO & Exportkontrolle
- DSGVO: Bleiben die Daten auf eurer Hardware, gibt es keine Drittland-Übermittlung, keinen komplizierten Auftragsverarbeitungsvertrag mit einem US-Hyperscaler, kein Schrems-II-Problem. Für viele ist das der kürzeste Weg zur Rechtssicherheit.
- Exportkontrolle / Dual-Use: Für Zulieferer im Defence- und Marine-Umfeld sind Konstruktions- und Prüfdaten häufig exportkontrollrelevant und dürfen nicht ungefragt durch fremde Rechenzentren fließen. Lokale Verarbeitung hält sie im kontrollierten Bereich. Die konkrete Einordnung im Einzelfall gehört zu eurer Compliance — ich sorge dafür, dass die Technik ihr nicht im Weg steht.
Kontakt
Wenn ihr überlegt, ob sich ein eigener KI-Server lohnt: Schreibt mir eure Rahmenbedingungen, dann rechnen wir es durch.
E-Mail: markus.kremer@void-main.com
Mobil: +49 177 889 8179
Festnetz: +49 4941 9914271
Zurück zu KI-Beratung im Nordwesten · weiter zu KI-Implementierung und KI-Strategie
When on-premise makes sense
- The data can't leave. Contracts with automotive, marine or defence clients that rule out the cloud. Raw logs, design data, inspection records. When “send it to a US provider” isn't an option, the question isn't whether to go local but how.
- Continuous base load. If you infer around the clock, you pay for every token in the cloud. Your own hardware is almost nothing but electricity once it's bought.
- Latency and offline. Close to production, no dependency on the internet, no “your credit has run out”.
When it doesn't — the limits
This is the real selling point, because almost nobody says it out loud:
- Only occasional use? Then an API is cheaper than any server. A box that idles 95 % of the time doesn't pay off.
- You always need the single strongest frontier model? That doesn't run locally. Then the right path is a frontier provider (e.g. Anthropic Claude) — with an EU region and a data-processing agreement if needed.
- Highly variable load (nothing, then a huge peak)? The cloud scales that better than bought hardware that dozes most of the time.
- Nobody can run it? Without someone handling drivers, updates and monitoring, on-premise turns into a permanent building site. That's exactly where I come in.
On-premise isn't a creed. Often the right answer is a hybrid: the local model for sensitive raw data, a frontier model for the heavy reasoning — the very architecture running in production at CARIAD.
What it really costs — the Zuse example
Not a brochure, my own server: a used HP Z840 plus three RTX 3090s in an external rig — 72 GB of GPU memory at full bandwidth. Two cards run the language model (Qwen3.6-27B, 256K context, ~65 tokens/s), the third handles image generation. More on it in the HandsOn on Zuse.
Acquisition, used market — Zuse's real figures:
| Item | rough |
|---|---|
| 3× RTX 3090 (used, ~€900 each) | ~€2,700 |
| HP Z840 workstation (used, with 128 GB RAM) | ~€680 |
| External rig, riser cables, second GPU PSU | ~€330 |
| Total for 72 GB VRAM @ 936 GB/s | ~€3,710 |
You can start smaller: a used workstation plus one 3090 and an NVMe lands around €1,900–2,100 — enough for models up to ~14B comfortably, or the 27B with a trimmed context. 2026 context: the current memory crunch has pushed RAM, SSDs and GPUs up; a used 3090 now runs ~€1,000–1,200.
Electricity (assumption €0.30/kWh): a 3090 draws up to ~350 W, power-capped to ~250 W. Idle around the clock ~190 W, three cards under full load (capped) ~900 W. Realistically, depending on load, roughly €500–1,100 a year (about €40–95 a month).
Operation is the item brochures leave out: drivers, model updates, monitoring, the lessons that hurt (models belong on an SSD, Whisper hallucinates on silence). That costs no euros, but time — and it's why most people hand operation off.
Break-even: for a permanent base load, a server like this clearly undercuts any comparable long-term rental in the cloud. For a few requests a week, never. The decision runs somewhere between those two poles — and I work it out with you concretely, not with a brochure rule of thumb.
Choosing the model
What runs locally depends on the hardware. On two RTX 3090s (48 GB VRAM), a 27B model in int4 with a long context runs smoothly (~65 tok/s) — enough for RAG, extraction, agentic steps and classification. For the heaviest reasoning you reach for a frontier model. The real work is the split: what has to stay local, and what may — anonymised — go outside.
GDPR & export control
- GDPR: if the data stays on your hardware, there's no third-country transfer, no complicated data-processing agreement with a US hyperscaler, no Schrems II problem. For many, that's the shortest route to legal certainty.
- Export control / dual-use: for suppliers in the defence and marine space, design and inspection data is often export-control relevant and mustn't flow through third-party data centres unasked. Local processing keeps it inside the controlled perimeter. The specific classification in each case is part of your compliance — I make sure the technology doesn't get in its way.
Contact
If you're weighing up whether your own AI server is worth it: send me your parameters and we'll work it out.
Email: markus.kremer@void-main.com
Mobile: +49 177 889 8179
Landline: +49 4941 9914271
Back to AI consulting in the North-West · on to AI implementation and AI strategy