Choosing the parts
Then came the big question: what should run the whole thing? I set up Dify, a visual agent builder, repaired it enough to judge it fairly, and said no. In the roadmap's words, it “exited the primary roadmap instead of expanding into framework soup.” LangGraph became the code-first orchestrator, n8n took visual automation, and MCP, the open standard for exposing tools to AI models, became the boundary every tool crosses.
What it taught meNew technology has to solve a problem I actually have, not just be new.
Operating it
April looks quiet in a feature list and matters more in hindsight. The work was a reconciliation of every record against what was actually running, a one-click health report, settings backup with a tested restore, a known-good checkpoint, startup validation, a registry of everything Falkor can do, and a control center to run it all.
The questions had changed. Not “can the AI do this?” but: what is really running, can I restore it, can I find a capability, and does startup reproduce a known-good state? By the end of April, ten build phases were done.
What it taught meIf you can't observe it and restore it, it isn't a product yet.
Making it a product
May was the busiest month of the project, with 826 commits. On 12 May the status panel learned a principle I still use: offline is only fine if offline is what I wanted. Every service is now compared with its desired state, not just shown as a green or grey light.
On 13 May, OpenClaw, a self-hosted agent platform, became part of Falkor on Falkor's terms: it runs on Falkor's own models and memory, so there is still one memory and one authority. When it later kept dropping, the cause turned out to be lower down, in the layer that keeps the Linux side of the PC alive, which taught me to diagnose from the bottom of the stack up. The same week, a custodian agent, Hermes, started watching the services and repairing what it safely could.
On 17 May I benchmarked the model candidates properly, for plain answers, reasoning, vision and structured output. The most promising large model was bigger than the GPU's entire memory, so it could never load fully, and running it half-loaded made the assistant feel slow. A smaller model that stays loaded makes a better household assistant than a stronger one that doesn't.
On 20 May a large reliability pass came back NO-GO: the fixes were real, but my first real use of the interface still found problems. The next pass fixed them and earned a GO the same day. The lesson became a rule for the AI agents building Falkor: an embedded widget isn't a launcher, navigating somewhere isn't an action, and a dry run isn't a working control.
On 24 May the product-complete wave closed, and the record called it ready for a massive re-audit. That was the right phrasing: complete meant complete against that month's definition of done.
What it taught meAI agents need a definition of done, the same as any team.
One product, and saying no
On 15 June, Falkor 3.0 turned a pile of features into one product: one navigation, one model authority and production baselines. The model setup got simpler too. Instead of a rack of chat models for different roles, there is one chat model that I choose, with specialized models only where chat isn't the job.
On 17 June I evaluated Agent Zero, a capable autonomous agent, and parked it. Its open-ended code execution, full desktop control and memory of its own conflicted with Falkor's approval rules and its single memory. By then, capability alone was no longer a reason to adopt anything.
What it taught meRestraint is an architectural skill.