werk 12 · the engineering discipline under the models

AI engineering.

The engineering layer an AI system is actually made of — the gateway every model call passes through, the context and retrieval pipelines, the evaluation harness, the cost and latency budget, and the release path that lets a model change without the product changing underneath it.

FamilyScale with AICarried byAdvisoryCarried byInterimProven lines3 / 8Detailed mandates4IndustriesConsumer Finance, IT Services, Insurance, Recruitment Software, Telecommunications

Werk 12 — who carries it

arms and mandates

A model is the smallest replaceable part of an AI product. Everything that decides whether it can be shipped, changed or trusted is engineering: where the call is made from, what context it is given, what it is allowed to do with the answer, how a regression is caught before a user meets it, and what a change costs in money and milliseconds.

This werk builds that layer. It assumes the model will be swapped — for a cheaper one, a newer one, a self-hosted one — and makes that swap a routine release rather than a rewrite.

The offer

what is bought, and in what shape

An AI system that can be operated by an engineering team rather than babysat by its authors: one governed path for every model call, an evaluation gate in CI, and a cost and latency budget that holds under load.

Model gateway and provider abstraction

A single governed chokepoint for every model call — routing, retries, quotas, PII masking and lineage — with a CI gate that fails the build on any call made outside it.

Context and retrieval engineering

Chunking, embedding, indexing and retrieval built against the real corpus, with the retrieval quality measured rather than assumed.

Evaluation harness and release gate

Golden sets, rubrics and offline plus online evaluation wired into the pipeline, so a prompt or model change ships on evidence.

Cost, latency and reliability budget

Token and inference budgets per route, caching and model routing, with the fallbacks that keep the product answering when a provider does not.

AI engineering review2 to 4 weeks

Gateway, retrieval and evaluation audit with a remediation sequence.

AdvisoryFractional AI engineering authority

Reference architecture, provider arbitration, release gates.

Interim seat6 to 18 months, Chief Data & AI Officer or CTO

The AI platform built, gated and operated by the client team.

Governed model gatewayVector and graph retrievalEvaluation harnesses and golden setsObservability and lineage loggingCI release gates

Used in these offers

what you can actually buy

State of the proof

counted from the ledger below
8sub-capabilities
3proven · a published case carries a sourced figure
4held · carried by a named operator, no case published
1declared · in scope, no published proof today

Experiences that prove it

4 detailed mandates

Sub-capabilities and evidence

8 lines, each with its state
Model gateway and provider abstractionheldOn-Kare: every model call routed through a single governed provider, with a CI gate failing the build on any call made outside it.
Context and retrieval pipelinesheldRetrieval over a knowledge graph built from a 15,500-file engineering corpus.
Evaluation harness and release gatingheldEvaluation dimensions, rubrics and reference datasets specified per AI phase before implementation.
Data platforms behind AI workloadsprovenBNP Paribas — FLOAAXA
Model lifecycle and deploymentprovenBNP Paribas — FLOA
Conversational and assistant engineeringprovenOrange BusinessPeetchr
Cost, latency and model routingdeclared
AI observability and lineageheldOn-Kare: model, prompt hash and timestamp logged on every AI output for conformance audit.

Publishable figures

named, sourced, attributable
225+AI functions specified · On-Kare platformOn-Kare venture · platform scope, 2026
0Model calls outside the governed gatewayOn-Kare venture · CI gate, no baseline, no exemption list

Questions

answered, in the open
How is this different from the AI werk?
The AI werk decides which model belongs behind which decision and proves it is worth putting there. AI engineering builds the system that call lives in: the gateway, the retrieval path, the evaluation gate, the cost budget and the release mechanics. One chooses the intelligence, the other makes it operable by a team that did not write it.
Why insist on a single model gateway?
Because governance is only enforceable at a chokepoint. Masking, routing, quotas and lineage are each trivial to implement once and impossible to enforce across scattered call sites, so the gateway is the difference between a policy and a claim.
What state is the proof in?
The data-platform and lifecycle lines are proven in client mandates — BNP Paribas FLOA and AXA — and the conversational lines at Orange Business and Peetchr. The gateway, retrieval and evaluation lines currently live inside the On-Kare venture rather than a client case, so they read HELD rather than PROVEN.

Related werks

same family, shared proof
Discuss a ai engineering mandateHow Advisory runs itAll twelve werks