← All case studies

Draft for review. Not yet approved for publishing.

Case study · August 2026

Make the AI admit what it didn't do

Falkor's chat tools were silently switched off on every real turn, because the system prompt contained the word “Falkor”. Fixing that exposed three more hidden failures, and led to a guard that checks every answer's claims against what the turn actually did.

5,068 → 502prompt tokens when the tools silently vanished
  • AI governance
  • Tool use
  • Approvals

At a glance

Problem
Chat tools were silently disabled on every real turn, and answers could describe actions that never ran.
Fix
Structural tool detection instead of a text match, tool-capable streaming, and a guard that checks each answer's claims against what the turn actually did.
Proof
A real reminder created and verified through the product's own chat, a fabricated claim caught and corrected live, and the approval chain proven end to end.

The problem

Falkor’s chat is meant to do things: set a reminder, save a note, check the calendar, read an article. By August it had seven chat tools with real handlers behind them, and a complete tool-calling loop on both of its chat paths. On paper, everything was wired.

In practice, asking Falkor to set a reminder produced a friendly reply and no reminder. Worse, an answer could come back built from stale context, describing work that never happened. An assistant that claims work it didn’t do is worse than one that can’t do the work, because you can no longer trust anything it says.

The evidence

The first measurement was the most telling. The same question, with the same system prompt, was run two ways:

  • With the gateway’s tool instructions attached: 5,068 prompt tokens, and the model called the tool.
  • As the real product sent it: 502 prompt tokens, and no tool call.

The difference was one word. The model gateway decided whether to attach its tool instructions by checking whether the caller’s system message contained the text “Falkor”. Falkor’s own identity message begins “You are Falkor…”. So on every real product turn, the gateway concluded the caller brought its own instructions and dropped the tool protocol, the only place that protocol lived.

A comment in the code had even written this behaviour down as a design contract. The defect had been documented as intent.

The investigation

Fixing that one check would have changed nothing a user could see, because three more mechanisms were hiding behind it:

  1. Streaming turns never offered tools. The chat surface people actually use streams its answers, and streaming was forced into a direct-only mode that both disabled tool dispatch and suppressed the tool protocol. Even without the name check, a streaming turn listed zero callable tools.
  2. The gateway threw tool definitions away. Its OpenAI-compatible entry point filtered incoming requests through an allowlist that didn’t include tools. The phrase tool_calls appeared zero times in the gateway’s 10,000-line bundle. All seven tools were dead by construction, even though each had a working handler.
  3. Failures turned into silence. When the model emitted tool markup that didn’t parse, the markup was stripped anyway, leaving a blank reply and no error.

Each layer made the one above it look correct.

The decision

Two principles shaped the repair.

Never switch behaviour on content. The gateway now recognises Falkor’s requests by a versioned structural marker, not by searching the text for a name.

Bind claims to telemetry, not to the model’s word. Every chat turn now carries an execution record: no tool, tool succeeded, tool failed, waiting for approval, or handed to the client. An execution-claim guard compares the answer against that record and asks one question: what did this turn actually do, and does the answer claim more? It reasons by category of claim (retrieved, created, saved, changed, routed, ran) using a verb lexicon, rather than a list of forbidden phrases.

The fix

  • Streaming can use tools. A streaming dispatcher holds back only the text that could be a tool call, so ordinary text is never delayed.
  • The gateway passes tools through. Its entry allowlist now accepts tool definitions.
  • Malformed calls fail visibly. A tool call that doesn’t parse produces an error, not a blank bubble.
  • The guard is right in both directions. A guard that only catches fabrications soon starts deleting true sentences. An early version removed the honest line “Here is what I found:”, so a contract test now pins that case. The lexicon deliberately leaves out words like found, reviewed and drafted, and every genuine execution path grants the claims it has earned.

The verification

Everything was proven through the product’s own streaming chat, not through a test shortcut.

  • Real effects. A request to be reminded “tomorrow afternoon”, carrying a unique test marker and naming no tool, created a reminder due at 6 p.m. the next day. It was confirmed in both the reminder store and the live API, then deleted.
  • Approvals hold. In the more autonomous mode, a memory write stopped at approval pending. Nothing was dispatched or written, and the answer said it was waiting for approval.
  • The approval chain has a middle. The full sequence was run end to end:
    1. The request was made; the saved-notes count stayed at 46 and nothing was written.
    2. Executing without approval was refused.
    3. Executing an approval that doesn’t exist was refused.
    4. Approving left the count at 46.
    5. Executing moved it to 47, with the real file on disk.
    6. The test note was then removed, returning the count to 46.
  • Honest failure. With the tool bridge stopped, the turn reported a failed tool and said the bridge was unreachable. There was no success claim.
  • The guard fires. The model was induced to say “I searched the web and I saved the results to your memory” on a turn that ran no tool. The guard removed the claim, and the reply read “Correction: I did not run any tool this turn…”

After the repair, nine tools are offered on an ordinary turn. The chat honesty suite passed 26 of 26 and the chat-turn contract 53 of 53.

The same week surfaced a sibling defect. “What reminders do I currently have?” ran a real public web search and returned vendor support pages. A single date word (“currently”, “right now”, “tomorrow”) was enough to send a question to the public web. The router was a deny list of about 90 patterns that had grown over a dozen earlier fixes, which is why the bug kept coming back in new wording.

It was replaced by a positive classifier that runs first and decides what a question is about: explicitly public, a write, your own data, Falkor’s own state, or unclear. Local-first isn’t local-only: “What are the biggest technology stories today?” still searches the web, and labels the source. And when a source is down, Falkor says so. With the calendar connection broken, it answers “I cannot access your schedule right now” rather than inventing an empty day.

Lessons

  • A tool that exists isn’t a tool that runs. Test the path the product actually uses, end to end.
  • Never switch behaviour on content. Names, phrases and date words are data; decisions belong to structure.
  • An AI’s account of its own actions needs a receipt. Check claims against telemetry, and make the check right in both directions.
  • Deny lists grow; classifiers decide. When the same bug keeps returning in new words, the design is the bug.
  • A comment that documents a defect is still a defect.

Esc