Case studies

Four engineering stories from Falkor

Each follows the same arc: problem, evidence, investigation, decision, fix, verification, lesson. Every number comes from Falkor's own dated records.

Execution-grounded claim verification

Make the AI admit what it didn't do

Published

Falkor's chat tools were silently switched off on every real turn, because the system prompt contained the word “Falkor”. The fix went further: each answer's claims are now checked against what the turn actually executed.

5,068 → 502prompt tokens when the tools silently vanished
  • AI governance
  • Tool use
  • Approvals

Preventing false-positive certification

Green is a claim, not a fact

Published

A certification harness that could declare success without evidence. A watchdog accusing a healthy store. 136 information sources shown as “active” that nothing ever ran. This is how Falkor learned to prove its own status.

6ways certification could pass without evidence, all closed
  • Quality engineering
  • Observability
  • Release governance

Governing AI-assisted software delivery

One builder, a fleet of AI agents

Published

Nearly 3,000 commits since March, most of them written by AI coding agents working in scoped lanes, with handoffs, evidence gates, independent review, and an operator who overturns a “PASS” that isn't one.

7 in 10main-repository commits signed by an AI coding agent
  • AI-assisted engineering
  • Program ownership
  • Agent orchestration

Stopping a rewrite and rebuilding in place

The rewrite that rebuilt Falkor

Published

A from-scratch replacement demoed with 0 of 13 user journeys working. It was frozen, audited for its best parts, and Falkor was rebuilt in layers instead.

184 → 98page routes in two days, with no capability lost
  • Program ownership
  • Architecture
  • Delivery judgment

Working on something similar? Discuss a project, or see how I work.

Esc