How I work
The operating model behind Falkor
AI agents implement scoped work. I own the architecture, the acceptance criteria, the evidence and the release decisions. It is how Falkor has been run since its first weeks, and it is how I work with teams.
- 01Plan
- 02Build
- 03Prove
- 04Ship
- 05Operate
The lifecycle
Seven stages, from plan to improve
I own Falkor's whole lifecycle: planning, architecture, delivery, verification, deployment, operations and audit. It has run with enterprise product disciplines from its first weeks: a roadmap, a groomed backlog, an agile cadence and a CI/CD pipeline.
01 · Plan
Plan it like a product
Falkor has had a phased roadmap since its first weeks: ten build phases by the end of April, then the Falkor 3.0 program. Work lives in a delivery tracker in which every item carries a phase, a gate, its dependencies, an owner and a completion check.
- Phased roadmap from month one
- Backlog captured and groomed as it grows
- Clear definition of done for every item
940
tracked items, 644 done
02 · Build
Build in scoped lanes
Work is cut into lanes: briefs with a goal, acceptance criteria, required reading, an ordered scope and a do-not-touch list. AI coding agents implement them. I own the architecture, the rules and the review.
- Change impact declared before any edit
- Several agents in parallel, with collision rules
- Every brief written to be checkable
541
named lanes in the canonical record
03 · Verify
Prove it before calling it done
Every new gate must be shown to fail before it is trusted to pass. Certification runs 56 gates over more than 7,000 browser tests, and a lane isn't done until its living documentation says so too.
- Negative controls on every gate: a test made to fail on purpose, proving the gate can catch a failure
- No retries or timeouts added to hide a red
- Behaviour suites for the product's promises
7,215 / 7,218
browser tests passed in one run
04 · Ship
Ship continuously, and safely
A self-hosted CI/CD pipeline carries every change. Blocking checks run first: change impact, documentation parity, audit and coverage. Deploys are atomic and refuse uncommitted code, the build ID comes from the commit, and a failed readiness check rolls the release back.
- Commit, built bundle and live build must match
- Automatic rollback on a failed readiness check
- Pre-commit hooks guard the history
Every deploy
commit, bundle and live build verified identical
05 · Operate
Run it every day
Watchdogs cover every core service, a proactive status check runs every 30 minutes, and readiness is checked after every boot. When something fails, it heals: the full stack comes back in 6.8 seconds.
- Status truth tested for false greens and false reds
- Game Mode and graceful restore
- Operator attests what only a person can check
11 / 11
critical readiness checks green
06 · Audit
Audit the whole thing
A read-only audit mapped 649 capabilities, 1,051 API operations and 7,331 dependencies, then triaged 412 findings by severity. It was checked for internal consistency before any of it entered the canonical docs.
- Every capability classified for privacy and cloud exposure
- Findings ranked and turned into roadmap items
- Documentation drift corrected in place
412
findings triaged by severity
07 · Improve
Feed it back
Findings become roadmap items for the next round. New open models, runtimes and open-source releases are evaluated as they land, and engine upgrades are re-validated against the full test battery before they stay.
- Audit-derived roadmap
- Upgrades validated, not assumed
- Retired features recorded, never silently dropped
0
test failures attributable to the 22 Sep engine upgrades
Operating rhythm
Every 30 minutes
A proactive status check, plus watchdogs on every core service.
Every day
Standup-style planning with the AI agents (what closed, what's next, what's blocked), defect triage, backlog grooming, and mixture-of-experts brainstorming: several AI models weigh in on open questions, then I decide and the decision is recorded. Software-update checks, memory re-indexing and fresh briefings run on their own.
Every week
A review of the AI landscape (new open models, runtimes and trending open-source projects), with mixture-of-experts sessions on the bigger decisions: what to adopt, what to retire, what comes next.
Every round
A planned round of parallel lanes with a steward, acceptance criteria and a definition of done. Nothing closes until its evidence and documentation are in.
In agile terms
| On Falkor | Agile equivalent |
|---|---|
| Round of parallel lanes | Sprint |
| Lane brief with checkable rules | User story with acceptance criteria |
| Delivery tracker | Product backlog |
| “Done-done” plus updated living docs | Definition of done |
| Daily planning and triage with the agents | Daily standup and bug triage |
| Mixture-of-experts brainstorming, decision recorded | Design review and decision record |
| Steward review | Tech-lead review and sign-off |
| Closeout: what's claimed, and what isn't | Sprint review |
| Audit-derived roadmap | Retrospective into backlog |
Working together
The same model, for your team
Advisory, hands-on or an audit: the same operating model, applied to your system.
- Advisory
- Reviews and working sessions with your team
- Hands-on
- Building alongside your team, from prototype to release
- Audits
- A fixed-scope assessment with a written report