werk 10 · agentic systems

Agentic.

Systems that act, not systems that answer — agents given tools, a boundary, an audit trail and an evaluation harness, so autonomy is engineered rather than hoped for.

FamilyScale with AICarried byAdvisoryCarried byInterimProven lines2 / 11Detailed mandates2SectorsTech & Telecom

Werk 10 — who carries it

arms and mandates

Agentic work starts where the demo stops. A model that answers well in a chat window is a capability; a system that books the appointment, files the claim or moves the record is an operator, and it inherits every obligation an operator has — a scope, a stop, a trace and a review.

The offer

what is bought, and in what shape

Autonomy that is engineered: agents with tools, a boundary, an audit trail and an evaluation harness, deployed behind a chokepoint you control.

Agent architecture and autonomy tiering

What the agent may do unaided, what needs a human, and the blast radius of each tier — written down before anything is wired to a tool.

AI gateway and tool integration

A single chokepoint for models and tools, MCP and function-calling integration, routing, quotas and cost control.

Guardrails and safety

Prompt-injection defence, PII masking, output validation and the escalation path when an agent hits its boundary.

Evaluation and observability

Golden sets, LLM-as-judge harnesses, trace-level lineage and the audit trail that makes an agent's action reconstructable after the fact.

Agent design sprint3 to 6 weeks

Agent architecture, autonomy tiers, guardrail and eval design.

AdvisoryFractional AI authority

Gateway architecture, safety review, vendor and cost arbitration.

Interim seat6 to 18 months, Chief Data & AI Officer or CTO

Agentic systems built, evaluated and operated.

MCPAI gatewaysKnowledge graphsEvaluation harnessesNeo4jModel routing

State of the proof

counted from the ledger below
11sub-capabilities
2proven · a published case carries a sourced figure
8held · carried by a named operator, no case published
1declared · in scope, no published proof today

Experiences that prove it

2 detailed mandates

Sub-capabilities and evidence

11 lines, each with its state
Conversational agents and dialog designprovenOrange BusinessPeetchr
Agent orchestration and multi-agent workflowsheldThe collective's own delivery runs on a supervised multi-agent pipeline; no client case published yet.
Tool use, function calling and MCP integrationheldModel Context Protocol servers and tool contracts wired across the collective's build stack.
Retrieval, memory and knowledge graphsheldKnowledge-graph retrieval built over a 15,500-file codebase corpus on the On-Kare venture.
AI gateway and single-chokepoint architectureheldOn-Kare: every model call routed through one governed provider, enforced by a hard CI gate with no exemption list.
Guardrails, prompt-injection defence and PII maskingheldOn-Kare: injection-pattern detection, PII masked before the model call, fail-closed consent guards.
Agent evaluation, golden sets and LLM-as-judgeheldEvaluation dimensions, rubrics and reference datasets specified per AI phase before implementation.
Lineage, observability and AI audit trailheldOn-Kare: model, prompt hash and timestamp logged on every AI output for conformance audit.
Human-in-the-loop and escalation designprovenOrange Business
Agentic FinOps and model routingdeclared
Autonomy tiering and blast-radius controlheldOn-Kare risk tiers T0 to T3: the lower the tier, the harder the stop — no auto-dismiss, dual control, hard blocks.

Publishable figures

named, sourced, attributable
149User paths shipped · 79 dialog flowsAnthony Chevalier · Lead Consultant Conversational AI, Orange Business (2020-21)
225+AI functions specified · On-Kare platformOn-Kare venture · platform scope, 2026
0Model calls outside the governed gatewayOn-Kare venture · CI gate, no baseline, no exemption list

Questions

answered, in the open
What does this werk actually build?
It builds the machinery around the model rather than the model itself: the tool contracts an agent is allowed to call, the gateway every call is forced through, the guardrails that run before the prompt leaves the building, the evaluation harness that decides whether the thing is fit to ship, and the audit trail that lets a regulator reconstruct a decision months later. The model is the smallest and most replaceable part of the system.
Why does the gateway matter more than the model?
Because governance is only enforceable at a chokepoint. Prompt-injection detection, PII masking before inference, cost routing and lineage logging are each trivial to implement once and impossible to enforce across forty scattered call sites. On the On-Kare platform every model call routes through a single governed provider, and a CI gate fails the build on any call made outside it — with no baseline and no exemption list, because the first exception normalises the second.
How is autonomy bounded?
By tiering it against blast radius before a line is written. A tier-zero route — one that can end a treatment, block a prescription or trigger an emergency — gets hard stops, dual control, a hold-to-confirm and no auto-dismiss. A tier-three route gets ordinary product ergonomics. The tier is declared in the file name, so a reviewer cannot miss which class of hardening applies.
Where is the proof today?
Two client mandates and one venture. Orange Business shipped 149 user paths including 79 dialog flows on a conversational platform, with the team scaling from 20 to 30 during delivery. Peetchr was built from scratch as an AI conversational-recruitment SaaS running on social messengers, carrying a French DeepTech label. The deeper agentic engineering — gateway, guardrails, evaluation harness, lineage — currently lives inside the On-Kare venture rather than in a client case, which is why most of this werk reads HELD rather than PROVEN.

Related werks

same family, shared proof
Discuss a agentic mandateHow Advisory runs itAll eleven werks