KI-Implementierung
AI implementation
Eine KI-Demo zu bauen ist heute leicht. Ein System zu bauen, das im Produktivbetrieb liefert, ist es nicht. Zwischen beidem liegt die eigentliche Arbeit — und die ist mein Geschäft.
Building an AI demo is easy today. Building a system that delivers in production isn't. The real work sits between the two — and that's my business.
Demo vs. Lieferung
Der Unterschied, ausformuliert, weil ihn kaum jemand ausspricht:
| Demo | Lieferung |
|---|---|
| Läuft einmal, im Meeting, auf dem Happy Path | Läuft unter echter Last, mit den hässlichen Eingaben |
| Kuratierte Beispiele | Fehlerbehandlung, Retries, Fallbacks, Guardrails |
| „Sieht gut aus“ | Messbar — Evals statt Bauchgefühl |
| Blackbox | Observability: man sieht, was es tut und was es kostet |
| Steht daneben | Integriert in Bestandssysteme (Tickets, DBs, APIs) |
| Der Erbauer muss es vorführen | Übergebbar: dokumentiert, wartbar, ein Team betreibt es |
Was ich baue
- Agentic Pipelines — mehrschrittige Agenten mit Tool-Use, die echte Arbeitsschritte erledigen, nicht nur chatten. Drill-Loops, die nachfassen, bis die Antwort trägt.
- RAG über euer Domänenwissen — Retrieval, das die richtigen Belege findet, mit Reranking. Kein Vektor-Voodoo, sondern nachvollziehbare Quellen.
- Tool-Use / Integration — das Modell ruft eure Systeme auf: Ticketsystem, Datenbank, interne API. Da, wo die Arbeit wirklich passiert.
- Hybrid lokal / Frontier, wo Datenhoheit es verlangt — die Aufteilung, die auf Souveräne KI on-premise beschrieben ist.
Woran man Lieferqualität erkennt
Fehlerfälle sind behandelt, nicht ignoriert. Es gibt Evals, mit denen man Änderungen misst. Man sieht Latenz und Kosten pro Vorgang. Es hängt an den Bestandssystemen, nicht daneben. Und jemand anderes als der Erbauer kann es betreiben. Fehlt eines davon, ist es eine Demo — egal wie gut sie aussieht.
Belege, keine Versprechen
- Agentic RCA-Pipeline bei CARIAD (produktiv): Hybrid aus lokalem Modell für Rohlogs und Frontier-Modell fürs Reasoning, agentischer Drill-Loop, RAG über kuratiertes Domänenwissen, Ticket-Twin-Detection, Vision für HMI-Screenshots. Validiert: 5,1 min/Ticket bei 0,66 $ — 70–180× ROI gegenüber manueller Analyse.
- LogAnalyzer — allein gebaut, mit Claude als Team an der Seite, von der ersten Codezeile bis in den Azure-Cluster. Der ganze Weg steht im HandsOn.
Kontakt
Wenn ihr einen Prototyp habt, der nicht über die Demo hinauskommt — oder gar keinen, aber ein echtes Problem: das ist der Punkt, an dem ich anfange.
E-Mail: markus.kremer@void-main.com
Mobil: +49 177 889 8179
Festnetz: +49 4941 9914271
Zurück zu KI-Beratung im Nordwesten · weiter zu Souveräne KI on-premise und KI-Strategie
Demo vs. delivery
The difference, spelled out, because almost nobody says it:
| Demo | Delivery |
|---|---|
| Runs once, in the meeting, on the happy path | Runs under real load, with the ugly inputs |
| Curated examples | Error handling, retries, fallbacks, guardrails |
| “Looks good” | Measurable — evals, not gut feeling |
| Black box | Observability: you see what it does and what it costs |
| Sits beside your systems | Integrated into existing systems (tickets, DBs, APIs) |
| The builder has to demo it | Handover-ready: documented, maintainable, a team runs it |
What I build
- Agentic pipelines — multi-step agents with tool-use that do real work, not just chat. Drill loops that keep digging until the answer holds.
- RAG over your domain knowledge — retrieval that finds the right evidence, with reranking. No vector voodoo, but traceable sources.
- Tool-use / integration — the model calls your systems: ticketing, database, internal API. Where the work actually happens.
- Hybrid local / frontier where data sovereignty demands it — the split described on Sovereign AI on-premise.
How to spot delivery quality
Error cases are handled, not ignored. There are evals to measure changes against. You can see latency and cost per operation. It hangs off your existing systems, not beside them. And someone other than the builder can run it. Miss one of those and it's a demo — however good it looks.
Evidence, not promises
- Agentic RCA pipeline at CARIAD (in production): a hybrid of a local model for raw logs and a frontier model for reasoning, an agentic drill loop, RAG over curated domain knowledge, ticket-twin detection, vision for HMI screenshots. Validated: 5.1 min/ticket at $0.66 — 70–180× ROI versus manual analysis.
- LogAnalyzer — built solo, with Claude as a team alongside, from the first line of code into the Azure cluster. The whole route is in the HandsOn.
Contact
If you have a prototype that can't get past the demo — or none at all, but a real problem: that's where I start.
Email: markus.kremer@void-main.com
Mobile: +49 177 889 8179
Landline: +49 4941 9914271
Back to AI consulting in the North-West · on to Sovereign AI on-premise and AI strategy