Capability audit · 25 Sep 2026

Your home.
Your data.
Your AI.

Falkor is a private, local-first AI system that coordinates our household's information, media, reminders and displays: chat, memory, news and automation, plus the operations that keep it running, all on one home PC under one set of rules. It's a personal system, not a product for sale.

Flight recorderReplays of recorded behaviour

How a request moves through FalkorEvery request passes Falkor's governed core, which routes it to tools, memory, local models or the collection engines, checks the answer against what actually ran, and returns it to you or your screens. An operations ring of watchdogs surrounds the core.Falkorgoverned coreYoubrowser · phoneToolsreminders · notesMemoryinbox → approvalModelslocal · one GPUEngines459 sourcesScreensTV · kiosks

You“Remind me tomorrow afternoon to call the plumber.”

  1. 01RequestRequest arrives from the browser
  2. 02RouteRouter: a write to your own data. It stays on the home PC
  3. 03Toolcreate_reminder · due tomorrow, 6:00 PM
  4. 04VerifyReminder store now holds the new item
  5. 05Guard“I've set a reminder” is backed by a successful tool call

FalkorDone. I'll remind you tomorrow at 6 PM.

Reconstructed from Falkor's logs and case studies; prompts are illustrative. Not a live connection.How the tools were fixed →

Pick a scenario to watch a real request move through Falkor, or hover a part of the map to see what it does.

What Falkor is

A household AI system, built and run like a product

Falkor began in March 2026 as a local AI stack and grew into one product: a single cockpit over local models that answer on the home PC, a long-term memory with review built in, engines that read hundreds of sources, and the operational tooling that shows what is actually working.

02

Governed actions

Consequential actions wait for approval, and each answer is checked against what the turn actually did, so the AI can't claim work it never ran.

Try the claim guard
03

Evidence over optimism

Status surfaces are tested for false greens as hard as for failures. One certification run covers 56 gates and more than 7,000 browser tests.

See the test results
04

Used every day

It drives the living-room TV, the morning news, the family calendar and reminders. It's daily infrastructure, not a demo.

See the real screens

The same patterns, in enterprise work

Falkor runs a household, but its controls are the ones teams need when AI touches real work.

In FalkorModels and memory on the home PC
At workKeeping data inside your own boundary, for regulated or sensitive work that can't leave it.
In FalkorApproval gates and write policies
At workA person approves consequential actions before an AI agent takes them, and the record shows what actually ran.
In FalkorWatchdogs and the custodian
At workOperations that repair what they safely can and hand the rest to a person, with the reason.
Bring them to your team

Everyday journeys

Three everyday journeys through the system

Simplified on purpose. The capability audit traces every path like these from the page you use, down through the APIs, libraries and services beneath it, to the data it touches.

Ask

A question becomes a governed turn on a local model.

  1. 01You askFrom the browser, a phone or the quick-ask drawer
  2. 02Governed turnOne local model you chose, a persona, local facts first
  3. 03Tools and memoryReminders, notes and recall; writes wait for approval
  4. 04Answer with receiptsClaims are checked against what actually ran

Public web search joins only when the question needs it.

Remember

Memory is earned: nothing becomes long-term without review.

  1. 01Capture“Remember this” in chat, the clipboard, or share to Falkor
  2. 02InboxEvery capture lands in a review inbox
  3. 03Your approvalPromotion to long-term memory is a decision, not a side effect
  4. 04RecallIndexed for search and used to ground answers

Memory stays on the home PC.

Stay informed

Hundreds of sources become one explained feed.

  1. 01459 sourcesNews feeds, APIs and sites, in one registry
  2. 02Two enginesFeeds and scraping. Duplicates are fetched once; failing sources are quarantined and re-probed
  3. 03One feedDe-duplicated, ranked and explained
  4. 04EverywhereNews, daily briefings, reports and chat answers

Source count as of September 2026.

Showcase

Twelve capabilities, up close

The four flagships first, then eight more. Each opens a real screen from Falkor, annotated, or a working demo you can drive yourself.

Live demo · sample turn

Chat

Remind me Friday to renew the car registration.

Done. I'll remind you Friday at 9 AM. I've also emailed you the renewal form.

Correction: no email was sent this turn.

Turn record

  • create_remindersucceeded
  • send_emailnever called

Claim guard: 1 unbacked claim struck and corrected.

The answer can only claim what the record shows.

AI · 01

Chat that can't pretend it acted

Ask in plain English for a reminder, a note or your schedule. Falkor does it with real tools, and every answer is checked against what the turn actually ran.

A fabricated claim was caught and corrected live

Read the case study: Make the AI admit what it didn't do

How it works

  1. Each turn leaves a record of every tool it called
  2. A claim guard compares the answer with that record
  3. Unbacked claims are struck and corrected before you see them
Live demo · sample captures

Review inbox · 3

  • From chat

    Prefers the local-AI briefing before sports.

  • From the share menu

    The car is due for service in May.

  • From the clipboard

    Recycling goes out Thursday night.

Long-term memory · 0

    Nothing yet. Only what you approve lands here.

    Nothing becomes memory without a decision.

    AI · 02

    Memory you approve

    Say “remember this” in chat, from the clipboard or from any app's share menu. It lands in a review inbox, and only what you approve becomes long-term memory that grounds future answers.

    Promotion is a decision, not a side effect

    How it works

    1. Captures land in a review inbox, never straight in memory
    2. Approved items are embedded locally and indexed for recall
    3. Rejected items never reach long-term memory
    Falkor's Self-Healing page: a clean bill of health with every required service healthy, and one optional integration idle by choice, with its cause, consequence and source explained.
    The Hermes custodian cockpit in Falkor: overall status green, last sweep 45 of 48 green and 0 red, and failure-ownership counters including 3,084 recovered and 0 missed.

    Operations · 03

    Heals itself, or says why not

    Watchdogs cover every core service and an always-on custodian sweeps every two minutes. Stop everything and the stack restores itself in 6.8 seconds; after a cold reboot it comes back unaided in about 14 minutes.

    Full stack restored in 6.8 s, measured 26 Aug 2026

    What you're looking at Self-healing

    1. You're asked to act only when automatic repair is impossible.
    2. Green only when every required service is healthy.
    3. An optional service idle by choice is explained, not painted red.

    What you're looking at Custodian cockpit

    1. Last sweep: 45 of 48 checks green, none red.
    2. A bounded sweep every two minutes, owned by the supervisor.
    3. Detected, delegated, recovered: 3,084 recoveries, 0 missed.
    Live demo · sizes illustrative

    One GPU

    1. Chat model resident, ready for questions

    One chat model you choose; scheduled work borrows the GPU and hands it back.

    Falkor's AI and Models page: one global chat model with a 32,768-token context, tool policy approvals required, VRAM free, and a persona using 296 of an 800-token prompt budget.
    Falkor's Model Maintenance page: the global model fully in VRAM, and a runtime residency panel comparing the model resolved for chat with the model actually active on the machine.

    AI · 04

    Local models that share one GPU

    You choose the chat model. Scheduled jobs such as the morning briefing load their own models on the same GPU and hand it back, and a guard refuses cloud models that pose as local ones.

    Chat, then a scheduled job, then chat again, each on the right model

    How it works

    1. One chat model, chosen by you, serves every surface
    2. Workloads borrow the GPU on schedule and give it back
    3. Residency is measured, never assumed from the selection

    What you're looking at AI & Models

    1. Tools wait for approval, by policy.
    2. Residency is checked live: the model is fully in GPU memory.
    3. The persona fits an 800-token prompt budget, 296 used.

    What you're looking at Model residency

    1. No fake “installed” or “ready” status.
    2. What chat asked for, next to what the machine actually has loaded.
    3. Memory pressure, measured live.
    The browser edition of Falkor's Local AI World report: a teal editorial masthead, then a section marked Partial whose shortfall note reads requested 10, delivered 9, 33 sources failed, no filler added.
    Falkor's Reports and Briefings page: a large serif masthead, report tabs from What Matters Today to Gadgets and Products, and cards for Local AI World and Geek Patrol, each marked published with quality 100.

    Data & integration · 05

    Briefings that write themselves

    Nine reports on the news, local AI, sports, entertainment and more come from one report engine, each with browser and email editions. When a report falls short, it says so instead of padding.

    Local AI World this week: 9 of 10 items delivered, 33 sources failed, no filler added

    What you're looking at Local AI World

    1. Every edition states its coverage: 46 of 51 items this week.
    2. Each section declares a target, a ceiling and a minimum it will accept.
    3. The shortfall is printed. Nothing is invented to fill the gap.

    What you're looking at Reports & Briefings

    1. All nine reports run on one canonical engine, one archive and one scheduler.
    2. News, local AI, sports, entertainment, charts and gadgets, one tab each.
    3. Each edition is scored before it publishes.
    A Falkor dossier on Ada Lovelace: a portrait, a summary, chapter navigation, and an at-a-glance panel listing 74 verified statements from 3 sources.

    Data & integration · 06

    Sourced dossiers

    Name a person, an organization, a product or a topic. Falkor researches official, reference and archival sources, then writes a sourced dossier from what it can verify, and shows what it found and what it couldn't.

    One dossier: 74 verified statements from 3 sources

    What you're looking at Sourced dossier

    1. Every statement opens its source.
    2. 74 verified statements from 3 sources, in edition 4.
    3. Chapters from the short version to the timeline, each searchable.
    Falkor's Spin Lens page: a dark banner titled Adversarial decision workbench, the selected local model and search setting, and a form to choose a proposition and its evidence.

    AI · 07

    An argument machine

    Spin Lens takes a claim, an article or pasted text and builds the strongest case for and against it, steelmans each side and pressure-tests the leader. It's a guard against one-sided answers, including the AI's own.

    Runs on the local model, with private search first

    What you're looking at Spin Lens

    1. Builds the case for and the case against, steelmans both, then rebuts them.
    2. Facts carry stable evidence IDs; assumptions and value judgments stay labeled.
    3. Nothing is saved unless you ask. Pasted text is never stored.
    Live demo · replays a measured run
    Sources
    16 / 110
    Stories
    —
    Duplicates
    0

    The recorded run: killed at source 16 of 110, then resumed with zero duplicate stories.

    Data & integration · 08

    Hundreds of sources, one clean feed

    459 sources across feeds, APIs and sites, collected by two engines. Duplicates are fetched once, failing sources are quarantined and re-probed, and a collection pass survives a crash.

    Force-killed at source 16 of 110, resumed with zero duplicate stories

    How it works

    1. Progress is checkpointed as the pass runs
    2. A resumed pass picks up where the last one stopped
    3. Stories are de-duplicated, so a restart can't count anything twice
    Falkor's Capabilities page, its live registry of runnable actions: 307 split into Can do now (107), Needs your OK (86) and Can't right now (114), with filter chips and cards showing health, write policy and locality.

    Operations · 09

    It knows what it can do

    One live registry, built from each owner's own records, lists every runnable action (tools, skills, MCP servers, scripts and recipes), whether it works right now and how it is used safely. The shelf explains; it executes nothing. It counts something different from the audit's 649 capabilities, which cover the whole stack.

    307 runnable actions in the live registry: 107 ready, 86 waiting for approval, 114 blocked with a reason

    What you're looking at Capability registry

    1. Ready now, waiting for your OK, or blocked, each with a reason.
    2. 566 source rows left out, each with a written reason.
    3. Every capability shows its health, its write policy and where it runs.
    Falkor's Open-Source Engines page: 7 of 7 engines online, 0 down, exposure loopback-only, and an evaluation confirming every running engine can be checked from Falkor without its own login.

    Data & integration · 10

    The best of open source, governed

    More than 20 open-source projects run as managed engines: photo library, document archive, private search, uptime monitoring, notifications, recipes, web archiving, workflow automation, design and vector search.

    • Immich
    • Paperless-ngx
    • SearXNG
    • Uptime Kuma
    • ntfy
    • Mealie
    • ArchiveBox
    • n8n
    • Penpot
    • Qdrant
    • ComfyUI

    Engine upgrades are re-validated against the full test battery

    What you're looking at Open-source engines

    1. Seven engines, all online.
    2. Loopback only: reachable from the home PC and nowhere else.
    3. Each engine is checked without its own login. Credentials are referenced by name, never by value.
    Live demo · sample services
    • Chat modelResident
    • Scheduled AI jobsOn schedule
    • Image engineReady
    • News collectorsNormal

    GPU

    1. Normal mode: Falkor's AI work holds the GPU

    Falkor holds the GPU for its AI work.

    Operations · 11

    Game Mode

    One switch hands the GPU to a game: background AI work pauses, services throttle, and everything is restored afterwards.

    Throttle and restore verified live

    How it works

    1. One switch, in Falkor's header
    2. AI work pauses, services throttle and the GPU is freed
    3. Switching back restores every service, and checks it
    Live demo · sample processes
    ProcessOwnerIdleAction
    falkor-webFalkoractive
    helper.exenone (orphan)14 min
    updater.exeunknown3 min
    worker.exenone (orphan)4 min

    Every stop passes four gates, and anything unknown fails closed

    1. Fresh census
    2. Verified orphan
    3. Idle 10 minutes or more
    4. Typed confirmation

    Pick a process and try to stop it.

    Operations · 12

    A process census with a safety catch

    See every Windows process, who owns it and whether it is idle. Stopping one takes a fresh census, a verified orphan at least ten minutes old and a typed confirmation, and every action is audited.

    Anything unknown fails closed

    How it works

    1. A fresh census before any stop
    2. Only verified orphans, idle ten minutes or more, qualify
    3. A typed confirmation, then an audit record

    Screens captured from the running system on 28 Sep 2026 and cropped to the page. Demos run in your browser on sample data.

    What it does

    Capabilities by outcome

    What Falkor does for the household, then how much of it is ready to run right now.

    • Ask

      Chat with local models that can set reminders, take notes and check the schedule, with every answer checked against what ran.

    • Remember

      Capture from chat, the clipboard or any share menu. Approved items become memory that grounds future answers.

    • Stay informed

      Hundreds of sources become one ranked, de-duplicated feed, nine reports and a daily briefing.

    • Create

      Local image and video generation, a media studio, and sourced dossiers on a person, organization, product or topic.

    • Automate

      Scheduled jobs, notifications and agents, with approval gates on consequential actions.

    • Operate

      Health checks, watchdogs, self-healing, certification and controlled deployment.

    307 runnable actions in the live registry

    The runnable actions in Falkor's live capability registry (tools, skills, MCP servers, scripts and recipes), each with live health and a safety gate. As of 28 Sep 2026.

    • 107 can run now
    • 86 need your OK first
    • 114 can't run right now, and say why
    The full audit: 649 capabilities in ten areasTechnical appendix · Capability audit, 25 Sep 2026

    How it counts: one capability is one thing Falkor can do, traced to its code, its screens and its API routes. The audit found 649 across 26 families, grouped here into ten areas. Each bar shows the share observed running during the audit (459 in all); a read-only audit can't exercise everything, so "not observed" doesn't mean broken. The registry above counts something narrower: actions that can be run on request.

    • Intelligence

      Local chat models, long-term memory with retrieval, personas and screen understanding.

      66 of 112

    • Information

      News from hundreds of sources, weather and radar, sports and local events: ranked, de-duplicated and explained.

      85 of 94

    • Life & household

      Reminders, the family calendar, documents, read-only email and shared household tools.

      79 of 93

    • Media & creative

      A living-room TV experience, music, radio and podcasts, video discovery, and local image and video generation.

      53 of 66

    • Home & displays

      Any screen can become a Falkor display: TV dashboards, weather radar and kiosks.

      Smart-home control is built, but the audit did not observe it running.

      7 of 14

    • Voice

      Push-to-talk speech in, natural speech out.

      Always-on listening is built and deliberately switched off.

      5 of 11

    • Automation & tools

      Scheduled jobs, notifications, a tool shelf and agents, with approval gates on consequential actions.

      54 of 68

    • Platform & side projects

      A registry and SDK that let new apps borrow Falkor's models, memory and status, plus a Labs shelf for side projects.

      66 of 79

    • Security & governance

      Origin guards, approvals, a local-only model rule, and every capability classified by privacy and cloud exposure.

      7 of 17

    • Operations

      Health checks, watchdogs, self-healing, certification and controlled deployment.

      A read-only audit can't exercise most of these, so many show as not observed.

      37 of 95

    Capability audit · 25 Sep 2026

    By the numbers

    System scale and reliability

    Three numbers first, each with what it means, then the supporting figures. Each is stated once, with its source and date, and none update themselves: the site changes only when a new snapshot is published.

    Breadth

    649 capabilities catalogued

    Everything Falkor can do, each traced to its code, its screens and its API routes. 459 were observed running during the read-only audit.

    Capability audit · 25 Sep 2026

    Verification

    7,215 of 7,218 browser tests passed in one certification run

    One run across 56 gates. The three failures were traced to their causes and fixed, and each passed when re-run on its own. There has been no full run since; the next one waits on a planned reboot and a set of manual checks.

    Certification · 22 Sep 2026

    Recovery

    6.8 s to restore the full stack

    With every service stopped at once, everything was back in 6.8 seconds. After a cold reboot, it recovers unaided in about 14 minutes.

    Recovery test · 26 Aug 2026

    Supporting figures

    Unless marked · Capability audit · 25 Sep 2026

    1,051

    API operations

    The endpoints the pages and services call

    100

    Pages in the cockpit

    Screens you can open

    7,331

    Mapped dependencies

    Links from pages to APIs, libraries, services and data

    381

    Tools

    Catalogued tools the system can call

    69

    Containers

    54 running at audit time

    36

    Scheduled background jobs

    Work Falkor runs on its own timetable

    78%

    Capabilities that are local-only

    504 of 649 never need the internet

    2,947

    Commits since March 2026

    Falkor and its stack

    Git · 28 Sep 2026

    AI practice

    How the AI architecture evolved

    Falkor has kept pace with a fast-moving field. Newer techniques joined the older ones rather than replacing them, and each was measured before it stayed. AI is a powerful tool with known failure modes, so the design works around what it gets wrong.

    What was added, and what it looks like in Falkor

    1. Context engineeringalongside prompt engineering

      Each turn's context is assembled from structure (the capability map, the persona, local facts and recalled memory), refreshed every 15 minutes and trimmed by policy.

    2. Retrieval-augmented generation (RAG)alongside keyword search

      Approved memories are embedded by a local model, indexed in a vector store and re-indexed nightly, so answers can find approved material before they are written.

    3. MCP toolsalongside workflow automations

      Capabilities are served as tools over the Model Context Protocol, a standard way to expose tools to AI models, from a universe of 462 tools across 8 providers, while n8n workflows stay behind approvals.

    4. Agentsalongside chat

      OpenClaw runs on Falkor's own models and memory, Hermes works as a 24/7 custodian that can repair services, and the agentic modes ask for approval before they act.

    5. Deliberate token budgetsalongside bigger context windows

      Budgets are measured per model: a reasoning model given 900 tokens produced nothing, and 2,400 produced a full answer.

    6. A model planealongside one big model

      A scheduler decides which AI model holds the GPU: one chat model you choose, workload models taking turns, and cloud models refused at the boundary.

    Where AI falls short, and the design answer

    Where AI falls shortWhat Falkor does about it
    Models claim work they didn't doEvery answer is checked against the turn's execution record.
    One date word can send a private question to the public webA positive classifier decides whose data a question is about before anything is searched.
    Reasoning models can spend the whole budget thinkingToken budgets are sized for thought and verified per model.
    A local model can crash on a large promptThe failure is reported and retried within bounds. The gap is never filled with invented text.
    AI coding agents over-report successGates that check the work, independent review, and an operator who signs off.
    GPU memory is finiteA model plane schedules who is resident; workloads hand the GPU back, and Game Mode frees it entirely.

    Reliability lessons from production use

    • Make the gate audit itself

      The rule that forbids hard-coded model names had exempted its own file, and was hiding two violations. It now scans its own source like everything else, and a missing self-scan fails the run.

    • Fetch once, account everywhere

      41 endpoints were shared by 157 sources. The first row fetches and every twin records the same result: 29.7% fewer wasted fetches, and no source hidden.

    • Quarantine, not retirement

      Sources that have failed 800 or more times in a row move to a 6 to 24 hour re-probe, and rejoin the moment one fetch succeeds. Nothing is deleted to make a dashboard greener.

    • Maintenance without downtime

      The engines can quiesce: hold new work, drain, checkpoint and keep running. A retention job deleted 871 rows in the middle of a collection pass without stopping it.

    • A flight recorder for every turn

      Each chat turn leaves a record of what actually ran. That record is how false claims get caught, and it's what the panel at the top of this page replays.

    The story

    How Falkor evolved

    Seven months, from a local AI stack in March to the full audit in September, told from Falkor's own records.

    1. Mar 2026

      A local AI stack gets a face

      Local models, memory and services on one home PC, then Falkor itself: one cockpit over everything the stack can do.

      Project records · Git history

    2. Mar to May 2026

      Ten build phases, then the expansion

      Memory, agent orchestration, source quality and integrations, built phase by phase. Then 826 commits in May alone: a media studio, a pixel-display companion and screen understanding.

      Roadmap · Git history

    3. 15 Jun 2026

      Falkor 3.0: one product

      From a pile of features to one experience with one governance layer: a single navigation, one model authority, and production baselines.

      Project records

    4. Jul to Sep 2026

      The featured story

      The rewrite that rebuilt Falkor

      Two AI models each designed a clean-sheet successor to Falkor. The winner, AURYN, then demoed with 0 of 13 user journeys working. I froze it, had it audited for its best coded parts and unbuilt ideas, and rebuilt Falkor one group of pages at a time. In the first two days, 184 page routes became 98 with no capability lost.

      Why it matters: stopping a rewrite that wasn't working, keeping its best ideas, and improving the working system in place.

      Read the case study
    5. Aug to Sep 2026

      Proving it

      Recovery programs, a 7,000-test browser battery made runnable, and certification gates that refuse to pass without evidence.

      Certification records

    6. 25 Sep 2026

      The full audit

      Every page traced down through the APIs, libraries and services beneath it to the data it touches. It's the snapshot this site is built from.

      Capability audit

    See the full timeline and the commit history

    Built by Dustin M. Gordon

    The same disciplines, for your team

    Falkor is where I practice them. The case studies tell three of its engineering stories in full, and the operating model shows how the work is run.

    Esc