Dustin M. GordonQuality engineering and AI modernization leaderProduct deliveryAI-assisted engineeringMission-critical systems

I build AI systems that can prove what they did.

I design the execution records, approval gates and quality checks that tie an AI system's claims to what actually happened. Falkor, the private AI system I built and operate for my household, is the working demonstration.

  • 20+ years in mission-critical software
  • Four years supporting IRS tax-processing modernization
  • Builder and operator of Falkor
Turn recordRecorded test · Aug 2026
The model drafted
(removed) “I searched the web and I saved the results to your memory.”
Tools the turn ran
None
Claim guard
No search or save is in the record, so the claim was removed
What you saw
“Correction: I did not run any tool this turn…”
A recorded test from Falkor. The answer can only claim what the record shows. See how it works

Flagship proofs

Four proofs you can inspect

Each pairs a problem with how Falkor handles it, a tested result and what the same pattern does at work. The demos run in your browser on sample data; the screen is real.

The disciplines behind them

Product ownership
a phased roadmap and a 940-item delivery tracker
Systems architecture
a whole-stack audit, traced from page to data
AI-assisted delivery
agent work briefs with review gates
Release engineering
atomic deploys that roll back on their own
Quality engineering
negative controls: gates shown to fail before they're trusted
AI governance
approval gates and a claim guard
Data reliability
crash-safe collection over 459 sources
Operations
watchdogs, recovery and self-healing
Technology adoption
20+ open-source projects, kept current
  1. RequestFrom a browser, a phone or the TV
  2. RouteDecides whose data it's about
  3. ToolsRun only when the turn needs them
  4. VerifyChecks what each tool changed
  5. GuardRemoves claims the record can't back
Tools and their checks run only when a turn uses a tool; every answer passes the guard.
  1. Live demo · sample turn

    Chat

    Remind me Friday to renew the car registration.

    Done. I'll remind you Friday at 9 AM. I've also emailed you the renewal form.

    Correction: no email was sent this turn.

    Turn record

    • create_remindersucceeded
    • send_emailnever called

    Claim guard: 1 unbacked claim struck and corrected.

    The answer can only claim what the record shows.

    01

    Chat that can't pretend it acted

    Problem
    Language models can describe work they never did.
    How it works
    Every chat turn keeps a record of the tools it ran, and a claim guard removes any claim in the answer that the record doesn't support.
    Result
    In the recorded test at the top of this page, the guard removed a claimed web search and memory save from a turn that ran no tool.
    At work
    The same check stops an assistant from reporting tickets, emails or deployments that never happened.
    See how claim verification works
  2. Live demo · sample captures

    Review inbox · 3

    • From chat

      Prefers the local-AI briefing before sports.

    • From the share menu

      The car is due for service in May.

    • From the clipboard

      Recycling goes out Thursday night.

    Long-term memory · 0

      Nothing yet. Only what you approve lands here.

      Nothing becomes memory without a decision.

      02

      Memory you approve

      Problem
      An assistant that saves everything also saves its mistakes, and then repeats them.
      How it works
      Anything you ask Falkor to remember lands in a review inbox first. Approved items are indexed on the home PC and ground future answers.
      Result
      Chat captures become durable memory only when you approve them. The audit also found side doors that skip review (agent tools that write notes directly, and rejected notes kept as reference) and lists them as open gaps.
      At work
      The same review step keeps a team's assistant from learning errors, or details nobody approved, as fact.
    • The Hermes custodian cockpit in Falkor: overall status green, last sweep 45 of 48 green and 0 red, and failure-ownership counters including 3,084 recovered and 0 missed.

      03

      Heals itself, or says why not

      Problem
      Home services fail when nobody is watching.
      How it works
      Watchdogs cover every core service, and a custodian agent sweeps every two minutes, repairing what it safely can and asking a person only when it can't.
      Result
      With every service stopped at once, the full stack was back in 6.8 seconds (measured 26 Aug 2026). By 28 Sep the custodian had logged 3,084 recoveries and missed none.
      At work
      The same loop keeps internal services up overnight and hands a person only what it can't fix.
    • Live demo · sizes illustrative

      One GPU

      1. Chat model resident, ready for questions

      One chat model you choose; scheduled work borrows the GPU and hands it back.

      04

      One GPU, shared on schedule

      Problem
      One home GPU has to serve chat, the morning briefing and image work without them colliding.
      How it works
      A scheduler decides which model holds the GPU. Jobs borrow it and hand it back, and what's loaded is measured, not assumed.
      Result
      Chat, then a scheduled job, then chat again, each verified on the right model. A guard refuses cloud models posing as local ones.
      At work
      The same scheduling lets teams share scarce GPUs and checks that every job ran on the model it asked for.

    See all twelve capabilities on the Falkor tour

    Case studies

    Four engineering case studies

    Each follows the same arc: problem, evidence, investigation, decision, fix, verification, lesson. Every number comes from the record.

    Execution-grounded claim verification

    Make the AI admit what it didn't do

    Falkor's chat tools were silently switched off on every real turn, because the system prompt contained the word “Falkor”. The fix went further: each answer's claims are now checked against what the turn actually executed.

    5,068 → 502prompt tokens when the tools silently vanished

    Preventing false-positive certification

    Green is a claim, not a fact

    A certification harness that could declare success without evidence. A watchdog accusing a healthy store. 136 information sources shown as “active” that nothing ever ran. This is how Falkor learned to prove its own status.

    6ways certification could pass without evidence, all closed

    Governing AI-assisted software delivery

    One builder, a fleet of AI agents

    Nearly 3,000 commits since March, most of them written by AI coding agents working in scoped lanes, with handoffs, evidence gates, independent review, and an operator who overturns a “PASS” that isn't one.

    7 in 10main-repository commits signed by an AI coding agent

    Stopping a rewrite and rebuilding in place

    The rewrite that rebuilt Falkor

    A from-scratch replacement demoed with 0 of 13 user journeys working. It was frozen, audited for its best parts, and Falkor was rebuilt in layers instead.

    184 → 98page routes in two days, with no capability lost

    All case studies

    The full tour

    Want to see the whole system?

    The full tour is for technical readers: the flight recorder, real screens, twelve capabilities up close, the capability audit, the architecture, measured results and the history.

    Enter Falkor

    How I work

    Plan, build, prove, ship, operate

    AI agents implement scoped work. I own the architecture, the acceptance criteria, the evidence and the release decisions.

    How work gets shipped

    1. Plan

      A phased roadmap, and a delivery tracker in which every item carries a gate, its dependencies and a check that shows it's done.

      Ten build phases delivered by the end of April, then the Falkor 3.0 program

    2. Build

      Work is cut into scoped briefs with acceptance criteria and a do-not-touch list. AI coding agents implement them; I review what they deliver.

      184 page routes rebuilt into 98, with no capability lost

    3. Prove

      Every gate is shown to fail before it is trusted to pass. Nothing is called done until it's documented and shown to work.

      Six ways certification could pass without evidence, found and closed

    4. Ship

      Gated, atomic deploys. The live build must match the commit, and a failed readiness check rolls the release back on its own.

      Every deploy: commit, bundle and live build verified identical

    5. Operate

      Watchdogs, a status check every 30 minutes and self-healing keep it running. Audits turn what they find into the next round of the roadmap.

      11 of 11 critical readiness checks green (Sep 2026)

    See the full operating model

    How decisions get made

    • Decision

      Red tests don't get longer timeouts

      Four of 56 gates went red in a 183-minute run. Longer timeouts were rejected; every failing check was fixed at its cause the same day.

    • Lesson

      A green gate that couldn't see

      A rule against hard-coded model names passed with 1,592 of them in the code. Its replacement now scans itself, and a missing self-scan fails the run.

    • Blocker

      The payout model broke the rules

      Paying donations straight to teachers would break district policy and gift limits, so the real-money pilot waits for a compliant payout model.

    Open the logbook: 29 entries

    About

    Built by Dustin M. Gordon

    I lead quality engineering and modernization for mission-critical software, with more than 20 years of delivery.

    Falkor is where I practice the whole lifecycle myself, from the roadmap to day-to-day operations. The habits are the ones I bring to every team: evidence before claims, gates shown to catch failure before they're trusted, and nothing called done until it works.

    What I bring

    • Product ownership from first idea to daily operations
    • AI-assisted delivery at scale: governed agents, scoped work, review gates
    • Quality engineering, certification and release governance
    • Private, local-first AI architecture
    • Fast, safe adoption of new AI patterns, models and open source
    • Plain reporting: what's done, what isn't, and why

    Beyond the résumé

    Online before the web
    I ran a BBS (a bulletin board system) for a couple of years, before the web took off, when a local dial-up modem was how you got online. Running a system out of my house clearly stuck.
    Gadget geek
    A/V guru and early adopter of new tech. Falkor driving the living-room TV is no accident.
    Virginia, through and through
    Born in Richmond, raised in Herndon, and I've lived all over Northern Virginia since, with family across the state.
    Dad
    My son is in elementary school.
    JMU alum
    James Madison University, where I studied computer information systems.
    Movies and TV
    A big fan of both, which may explain the A/V habit.
    DC sports fan
    And I grew up playing soccer and basketball.
    Percussion
    Grew up playing it, too.

    Work with me

    Turning AI experiments into dependable systems

    I'm open to full-time roles and to consulting. Here are four ways I can help, and how the work runs.

    • AI architecture review

      A structured look at an AI system or plan: where it can fail, what it claims versus what it can show, and a ranked list of fixes.

    • AI governance design

      Controls that let AI act safely: approval gates for consequential actions, answers checked against what actually ran, and audit trails people can read.

    • Quality strategy and release gates

      A test strategy, certification gates that are shown to catch failure before anyone trusts them, and release criteria that won't pass without evidence.

    • Private and local AI prototypes

      A working prototype that keeps your data on your hardware: local models, retrieval over your own documents, an evaluation set and plain status notes.

    How engagements run

    Advisory
    Reviews and working sessions with your team
    Hands-on
    Building alongside your team, from prototype to release
    Audits
    A fixed-scope assessment with a written report

    Also available: delivery-process improvement for AI-assisted teams, and fractional AI architecture.

    Hiring for a full-time role

    Quality engineering, AI systems and product delivery, backed by more than 20 years of mission-critical work.

    Have a project

    An AI system that works in demos but is hard to trust in production? Send a short brief: what you're building, what's at stake, and when.

    Both open an email with short prompts. Or write to hello@dustinmgordon.com

    Prefer LinkedIn? Message me there (opens in a new tab)

    Copies the email address

    Esc