The journey

Building a local AI until it became a system

Falkor started as an experiment: could I run a good AI assistant on my own PC? Every time it grew, something failed in a way I hadn't expected, and each failure forced a sharper definition of what “working” means. This is that story, in five eras, from February to September 2026.

February to September 2026 · 9 min read

The dated record behind it is the logbook, and the engineering stories are told in full in the case studies.

Illustration: a glowing path through a night landscape to a house on a hill, with five stations: foundation, assistant, platform, truth and governance.(opens the full-size image)
AI illustration The five eras, from a local-first experiment to a governed home system.

Era 1February 2026

Foundation

It started with a purchase. In February I bought a Windows PC with one strong GPU, specifically to run AI at home, and spent the first weeks on research and the basics: Ollama to serve local models, Open WebUI to talk to them, and a lot of reading about what a private, always-available assistant would take.

Before there was much code, I wrote down the mission: a local-first assistant, private by design, with swappable brains. Beside it went a rule that sounded minor and became the most durable decision of the project: role names stay stable, and model names are replaceable. So did a warning to myself: avoid architecture churn and framework fashion.

Illustration: a glowing chip with a brain at its center, wired to three labels: local-first operation, privacy-respecting, swappable brains.
AI illustration The mission, as the first records put it.

What it taught meDecide what must never change before you choose the parts that will.

Era 2March 2026

Assistant

March is where the records begin. I started a checklist, a roadmap and handoff notes that every later change had to keep true, and I've kept them ever since.

Voice came first. Wake words, room behavior and playback went through several rounds before they felt dependable, and web search joined the voice path. Even the voice itself was chosen by measurement: a newer text-to-speech engine was slower for quick replies, so the original stayed.

Memory became a design rather than a folder. The first version removed duplicates, ranked what it kept and suggested what to promote, instead of turning every conversation into permanent memory.

What it taught meAn assistant becomes useful when it can listen, search and remember, deliberately.

Era 3Late March to June 2026

Platform

Choosing the parts

Then came the big question: what should run the whole thing? I set up Dify, a visual agent builder, repaired it enough to judge it fairly, and said no. In the roadmap's words, it “exited the primary roadmap instead of expanding into framework soup.” LangGraph became the code-first orchestrator, n8n took visual automation, and MCP, the open standard for exposing tools to AI models, became the boundary every tool crosses.

What it taught meNew technology has to solve a problem I actually have, not just be new.

Operating it

April looks quiet in a feature list and matters more in hindsight. The work was a reconciliation of every record against what was actually running, a one-click health report, settings backup with a tested restore, a known-good checkpoint, startup validation, a registry of everything Falkor can do, and a control center to run it all.

The questions had changed. Not “can the AI do this?” but: what is really running, can I restore it, can I find a capability, and does startup reproduce a known-good state? By the end of April, ten build phases were done.

What it taught meIf you can't observe it and restore it, it isn't a product yet.

Making it a product

May was the busiest month of the project, with 826 commits. On 12 May the status panel learned a principle I still use: offline is only fine if offline is what I wanted. Every service is now compared with its desired state, not just shown as a green or grey light.

On 13 May, OpenClaw, a self-hosted agent platform, became part of Falkor on Falkor's terms: it runs on Falkor's own models and memory, so there is still one memory and one authority. When it later kept dropping, the cause turned out to be lower down, in the layer that keeps the Linux side of the PC alive, which taught me to diagnose from the bottom of the stack up. The same week, a custodian agent, Hermes, started watching the services and repairing what it safely could.

On 17 May I benchmarked the model candidates properly, for plain answers, reasoning, vision and structured output. The most promising large model was bigger than the GPU's entire memory, so it could never load fully, and running it half-loaded made the assistant feel slow. A smaller model that stays loaded makes a better household assistant than a stronger one that doesn't.

On 20 May a large reliability pass came back NO-GO: the fixes were real, but my first real use of the interface still found problems. The next pass fixed them and earned a GO the same day. The lesson became a rule for the AI agents building Falkor: an embedded widget isn't a launcher, navigating somewhere isn't an action, and a dry run isn't a working control.

On 24 May the product-complete wave closed, and the record called it ready for a massive re-audit. That was the right phrasing: complete meant complete against that month's definition of done.

What it taught meAI agents need a definition of done, the same as any team.

One product, and saying no

On 15 June, Falkor 3.0 turned a pile of features into one product: one navigation, one model authority and production baselines. The model setup got simpler too. Instead of a rack of chat models for different roles, there is one chat model that I choose, with specialized models only where chat isn't the job.

On 17 June I evaluated Agent Zero, a capable autonomous agent, and parked it. Its open-ended code execution, full desktop control and memory of its own conflicted with Falkor's approval rules and its single memory. By then, capability alone was no longer a reason to adopt anything.

What it taught meRestraint is an architectural skill.

Era 4June to August 2026

Truth

Audits, and a rewrite put to the test

A hard audit on 3 June came back with serious findings, and a focused repair lane cleared them. A backup audit then found the quietest kind of gap: newer apps kept their data in Docker volumes that the file-based backup checks couldn't see. Nothing had broken; it only looked backed up. Backup coverage now starts from an inventory of every place data lives.

In July I put the biggest question of all to a real test: would a clean start beat what I had? Two AI models each designed a complete successor, and the winner, AURYN, looked better on paper. In August I stopped it: its demo couldn't run the journeys that mattered. I kept its best ideas and rebuilt Falkor one layer at a time instead, and in two days 184 page routes became 98, with no capability lost.

What it taught meJudge by what actually works, not by what looks right on paper.

When execution isn't truth

On 21 August I asked Falkor what reminders I had. It ran a real web search and answered with vendor help pages. Nothing had been faked; the search really ran. It had asked the wrong source. Falkor now decides where the truth lives before it decides how to look.

Illustration: a glowing signpost at a crossroads of data streams, one arm pointing to local truth, the other to the public web.
AI illustration Where does the truth live? Decide that before deciding how to look.

The same week taught three more lessons. Two model settings agreed with each other while both ignored the model I had chosen, so agreement between derived values proves nothing until it's checked against the source. A test that stopped early looked like an honest failure while skipping every check after the stop, so a failing exit no longer counts as full coverage. And once the status screens were willing to show red, one of them turned out to recommend starting a service that could never start. It now says plainly that Falkor can't recover it on its own. In the record's words: “Making the red visible exposed a false instruction.”

Falkor's own scheduler also became the system of record for timed jobs. n8n was retired, later restored behind approvals, and whether it stays is still an open question.

What it taught meA system can execute perfectly and still answer wrongly if it asks the wrong source.

Era 5September 2026

Governance

September was less about new features than about removing the ways existing ones could mislead, collide or overstep.

Watchdogs learned to tell alive from working, because a process can be up while its scheduler has quietly stopped. Backups are judged by what is actually safe, not by whether a job ran. Memory follows the collection I've selected, and startup is proven by what comes up, not by what was launched.

One test meant for an isolated copy reached the live system. Now every test takes its target from the test runner, and a guard fails certification if one ever points anywhere else. Test isolation isn't a convention; it's a security property.

Behavior suites now judge Falkor's answers against a fixed baseline, not just status codes, and they found real answer defects on their first run. Every action that changes something is converging on one path: a policy check, approval where needed, the intent recorded, the effect, a check of the result and a receipt. And the oldest rule from February finally became enforceable: hard-coded models, engines and voices came out of the code that runs every day and into settings I choose, and a guard now refuses new ones.

Illustration: a vertical chain of six glowing blocks: request, policy, approval, intent, execution, and verification with a receipt.
AI illustration The path every action that changes something is converging on.

On 22 September a full certification run passed 7,215 of 7,218 browser tests with retries off, and on 25 September a whole-stack audit mapped 649 capabilities.

What it taught meAn AI saying “I did it” isn't a receipt.

Now30 September 2026

Where it stands

Today the deployed build matches its source exactly, and the chat model is a setting: Gemma 4 12B as of 30 September. That's today's choice, not Falkor's identity. The point of eight months of work is that it can change again without Falkor having to become something else.

Falkor began as an attempt to run a good AI on my own PC. It became an exercise in product architecture, systems integration, reliability, governance and keeping the records true, and in running AI as a real household utility.

What it taught meThe models and tools are replaceable; the authority, policy and evidence boundaries are not.

Models

Models that came and went

Dated snapshots, not rankings. The role stays; the model in it is a setting.

  • Piper and KokoroVoice

    What I learned
    The newer engine was slower for quick replies
    What I did
    Piper stayed for the fast voice
  • Gemma 4 26BChat

    What I learned
    Larger than the GPU's entire memory; half-loaded, it felt slow
    What I did
    Never the daily model
  • Hermes 3Chat

    What I learned
    Tested as a quick candidate: a model, not a framework
    What I did
    Stayed a swappable candidate
  • Qwen 2.5 Coder 14BChat

    What I learned
    Practical enough for a major checkpoint
    What I did
    The chat model at May's product-complete wave
  • Qwen 3.5 9BChat

    What I learned
    Fit the GPU and stayed loaded
    What I did
    The one chat model once the role rack was retired
  • LLaVA 13BVision

    What I learned
    Read a real screenshot within the machine's limits
    What I did
    Kept as the vision model
  • Gemma 4 12BChat

    What I learned
    Fits the GPU and stays loaded
    What I did
    The chat model as of 30 September 2026

Choices

What I adopted, and what I said no to

Every new tool had to solve a problem or replace something. Most of the value is in the no's.

Illustration: how the stack evolved in five layers: foundation, interaction, orchestration, platform and governance, each with its lesson, and Dify and Agent Zero marked as rejected or parked.(opens the full-size image)
AI illustration How the stack evolved, layer by layer. Select it for the full size.
  • DifyEvaluated, then declined

    Another orchestration layer wasn't worth it: no framework soup.

  • LangGraphAdopted

    Code-first orchestration for workflows that need state and approvals.

  • n8nAdopted, retired, restored

    Good for visual automation. Falkor's own scheduler now runs the timed jobs, and n8n's future is an open question.

  • MCPAdopted as the boundary

    One standard way to expose tools, with the rules kept in Falkor. It now serves 462 tools across 8 providers.

  • OpenClawIntegrated on Falkor's terms

    It runs on Falkor's own models and memory: one memory, one authority.

  • Agent ZeroParked

    Open-ended autonomy conflicts with Falkor's approval rules.

  • A chat model per roleRetired

    One chat model I choose, with specialized models only where chat isn't the job.

  • Cloud models as stand-insRefused

    Local-first is a boundary: a cloud model is never a silent substitute.

Lessons

Eight lessons

Illustration: six things that went wrong, each paired with the rule it created, under the line “A green process is not a green product.”(opens the full-size image)
AI illustration Failures that changed Falkor, and the rule each one created. Select it for the full size.
  1. Stable roles beat favorite models.Models changed too often to be the architecture. The roles stayed; the models rotated.
  2. Local-first is a policy, not a slogan.Falkor reads the web when a question needs it, but where the models run and who is in control are fixed boundaries.
  3. A green process isn't a green product.Alive, listening and healthy-looking aren't the same as working. Measure progress, not existence.
  4. Right execution can still give a wrong answer.Choosing the source is part of the architecture, not the prompt.
  5. AI action needs evidence outside the AI.“Done” needs a receipt: what was intended, what policy allowed, what ran and what changed.
  6. Tests need adversarial engineering too.A test can pass on nothing, stop early or reach the wrong system. Prove what it covered.
  7. New technology must displace or solve something.Dify declined, Agent Zero parked, OpenClaw bounded: selectivity kept Falkor from collapsing under its own experiments.
  8. “Complete” is a dated claim.Complete as of this evidence, never finished forever.

You choose the destination. Falkor helps you get there.

Written from Falkor's own dated records: the project checklist and roadmap, the handoff notes, the truth policy and the Git history. The records begin in March 2026; February is from my own memory. The illustrations were generated with AI (Google Gemini and ChatGPT).

Esc