Articles
Notes on architecture, AI agents and production LLM systems — in series and standalone posts.
Latest · 21
- 07 Video Generator, part 7: testing whether a character stays the same · Video Generator Two characters, a croissant, and experiments in Vini Studio. How I test character consistency, reproduce failures, and revise a hypothesis when the next test contradicts it.
- · How much of a company can one person hold in their head? AI lets one person do more and more. What will remain of familiar teams, and how many decisions can that person handle?
- 02 CRM on Markdown, part 2: 27 hours, 23 commits and a catalog of failures · CRM на Markdown The anatomy of a Markdown CRM built with an AI agent on a live project — told through its failures, because that's where every rule came from. Someone else's words recorded as the owner's voice, a ready-made comment posted under a stranger's post, a user's silence explained by the product while the user was moving house. Three breeds of failure — data, process, interpretation — each with its mechanism and price; four decisions I reversed within hours of making them; and the main lesson of the whole exercise: a rule without a guard does not survive, even for the person who wrote the rule two paragraphs above.
- 01 CRM on Markdown, part 1: how to create a CRM in an evening — step by step, with an AI agent that finds leads · CRM на Markdown A step-by-step guide to a working CRM without a database, SaaS or a single dependency: install Claude Cowork or Claude Code, add the md-crm skill in two commands or one ZIP, connect the Chrome extension — and the agent triages LinkedIn search results you scroll, files contact cards and keeps the board honest. You talk to it in plain words: add a contact, log her reply, what's on fire. Same skeleton runs an ATS for hiring. It ends with the final split — what the agent takes over and what stays human — and a way to reach me if you want it wired to your funnel.
- 11 My Architect, part 11: eight hundred commits I didn't write · My Architect Not a demo but a live project — a video pipeline with 802 commits in 78 days, with a front end, a back end, video rendering, Python sidecars and five model vendors, and my_architect holding it all together. Inside: the real numbers from the tracker (650 nodes, 296 requirements, 63 RFCs), how a requirement is wired to a commit and a test through a coverage gate that fails the build, why I draw HTML wireframes before any code, what commit timestamps actually measure in the story of “I hand a task over at night and by morning it's done” — the day has two humps with the night between them, and morning commits are twice the size of daytime ones — and why all of this opens a door for teams instead of closing it.
- 06 From canvas to studio: a reusable library, parallel runs and auto-packaged clips · Video Generator Earlier the canvas assembled a single clip. Now it has a name — Vini Studio — and it grew from a graph into a workspace where I can turn out clips in batches. Inside: why the element library and the render cache are one mechanism over an artifacts table; a prompt manager with system and personal templates; a node that compresses the previous episode into ~1000 tokens of memory; parallel runs and the topological-sort bug that made scenes start before the scriptwriter feeding them; subtitles built not from my script but from what was actually said on screen; a design system I generated with a single prompt, and packaging rendered in Remotion. Plus an honest look at the numbers — twelve dollars a clip and 22% full-watch retention.
- 05 Keeping the same characters across AI clips: universal elements instead of the first frame · Video Generator In episode two of "Sheriff Potato" the cast grew, and the first-frame trick stopped carrying it — to keep characters from drifting between scenes I moved the canvas to references, my "universal elements". Inside: why video models force you to choose between a first frame and elements (and why Kling is the exception), what 8 seconds cost on Grok, Gemini and Seedance, the 480p trick, ImageRef tags and characters' distinguishing marks. Plus an honest look at scene-six bugs — and the finished episode at the end.
- 04 My node canvas for video: building a whole clip on a single graph · Video Generator The first three parts took the pipeline apart piece by piece — the LLM for the scriptwriter, the video model, the storyboard. Here I show the whole canvas at once: my own node editor on litegraph and the ComfyUI core, where Claude writes the script, characters live as reusable "souls", and the "one frame → video" trick holds a scene's opening. Plus the first clip the canvas assembled on its own, and the "cabbage hat" bug that taught me the most.
- 10 My Architect, part 10: a repo memory for the agent — a code graph it's not allowed to trust · My Architect The recursive-context story continues: four releases in one day teach the agent a persistent code graph (Graphify) under strict distrust rules — freshness gate, facts only from live files. Plus a controlled A/B test of graph vs grep on NestJS: −71% tokens on subsystem understanding, parity elsewhere — and the hidden sub-agent costs that make naive benchmarks lie.
- 09 My Architect, part 9: Recursive context — teaching the agent to admit what it hasn't read · My Architect From MIT's Recursive Language Models paper to the recursive-context skill in one day: why a huge context window doesn't solve big data, how a fan-out of isolated sub-agents with honest coverage replaced a Python library, and what only live runs could catch — from a baseline that turned out too good to a 14-millisecond bug.
- 08 My Architect, part 8: Event Storming as sequence control, not another diagram · My Architect A picture-diagram checks nothing. Event Storming does: every event has a cause, every command an effect. How an agent runs a board through a deterministic sequence analyzer, closes mechanical gaps to zero, and leaves unresolved business questions visible. On a live Forklift project — 5 contexts, 17 gaps caught.
- 03 Storyboards before the video: I see the whole cut early — and fix the script while it's cheap · Video Generator Between the scriptwriter and the video model my pipeline got a layer that changed everything — the storyboard. A pencil storyboard shows the whole cut before I spend a second on video: I read it, catch a flat line or a lost character, and fix the script while it costs pennies. Inside — a four-stage evolution, my full prompts, and a bonus comparison of four models that draw the storyboard.
- 02 Grok Imagine 1.5: testing the new video model — experience and comparison · Video Generator Grok Imagine 1.5 is out — xAI's fresh video model with native audio and lip-sync. I ran it through my pipeline on a deliberately tough scene and compared it with Seedance 2.0, Kling v3 and Veo 3.1. Inside — four videos, my experience with Grok, and an honest breakdown: where it wins and where it doesn't.
- 01 Choosing an LLM for a production job: a blind, cross-vendor evaluation · Video Generator One LLM call sets the quality ceiling of every video my pipeline makes — and is paid on every one. So I benchmarked four models blind, across vendors, on the real production prompt. The newest model won the narrow metric and lost overall. The method and the quality-per-dollar decision.
- 07 My Architect, part 7: Obsidian, Claude Code docs and OpenClaw — field notes · My Architect I tried each of them for one reason: to be productive on real projects with AI. What job Obsidian, the Claude Code docs approach and OpenClaw actually do, where each hits its ceiling, and what that taught me about project memory.
- 06 My Architect, part 6: the human side — WBS, Building Blocks and the User Story Map · My Architect The series covered the agent; this part is about me. How an architect uses My Architect by hand: WBS mapped onto building blocks, a User Story Map for prioritization, and why a diagram that an agent executes beats any whiteboard.
- 05 My Architect, part 5: shipping skills and MCP via the Claude Code marketplace · My Architect An MCP server, a skill and four slash commands used to take a manual to install. The Claude Code plugin marketplace reduces it to two commands: how the manifest works, why the skill lives in a public repo, and what version discipline costs. Final post of the series.
- 04 My Architect, part 4: the agent's working loop · My Architect How an agent runs a project for weeks: get_next_task picks the work, docs stay current while the code is written, deferred items become nodes, and /reconcile keeps the plan honest against the codebase.
- 03 My Architect, part 3: a planning model that makes agents think like architects · My Architect How My Architect models a project so the agent plans like an architect: node hierarchies with presets, four requirement types inherited down the tree, a title linter and a User Story Map the agent reads as text. Third post in the series.
- 02 My Architect, part 2: an architecture with no database · My Architect Files are the agent-native choice, a database is dead weight here. YAML and Markdown as storage: atomic writes via rename, a per-project mutex after a real lost-update bug, live sync through chokidar and WebSocket.
- 01 My Architect, part 1: giving an AI agent memory and a plan · My Architect An AI agent forgets the project after every session. How My Architect keeps the plan, requirements and decisions alive between sessions: a UI for humans, MCP for agents. First post in a series of seven.
Series · 01–07
Video Generator
- 01 Choosing an LLM for a production job: a blind, cross-vendor evaluation One LLM call sets the quality ceiling of every video my pipeline makes — and is paid on every one. So I benchmarked four models blind, across vendors, on the real production prompt. The newest model won the narrow metric and lost overall. The method and the quality-per-dollar decision.
- 02 Grok Imagine 1.5: testing the new video model — experience and comparison Grok Imagine 1.5 is out — xAI's fresh video model with native audio and lip-sync. I ran it through my pipeline on a deliberately tough scene and compared it with Seedance 2.0, Kling v3 and Veo 3.1. Inside — four videos, my experience with Grok, and an honest breakdown: where it wins and where it doesn't.
- 03 Storyboards before the video: I see the whole cut early — and fix the script while it's cheap Between the scriptwriter and the video model my pipeline got a layer that changed everything — the storyboard. A pencil storyboard shows the whole cut before I spend a second on video: I read it, catch a flat line or a lost character, and fix the script while it costs pennies. Inside — a four-stage evolution, my full prompts, and a bonus comparison of four models that draw the storyboard.
- 04 My node canvas for video: building a whole clip on a single graph The first three parts took the pipeline apart piece by piece — the LLM for the scriptwriter, the video model, the storyboard. Here I show the whole canvas at once: my own node editor on litegraph and the ComfyUI core, where Claude writes the script, characters live as reusable "souls", and the "one frame → video" trick holds a scene's opening. Plus the first clip the canvas assembled on its own, and the "cabbage hat" bug that taught me the most.
- 05 Keeping the same characters across AI clips: universal elements instead of the first frame In episode two of "Sheriff Potato" the cast grew, and the first-frame trick stopped carrying it — to keep characters from drifting between scenes I moved the canvas to references, my "universal elements". Inside: why video models force you to choose between a first frame and elements (and why Kling is the exception), what 8 seconds cost on Grok, Gemini and Seedance, the 480p trick, ImageRef tags and characters' distinguishing marks. Plus an honest look at scene-six bugs — and the finished episode at the end.
- 06 From canvas to studio: a reusable library, parallel runs and auto-packaged clips Earlier the canvas assembled a single clip. Now it has a name — Vini Studio — and it grew from a graph into a workspace where I can turn out clips in batches. Inside: why the element library and the render cache are one mechanism over an artifacts table; a prompt manager with system and personal templates; a node that compresses the previous episode into ~1000 tokens of memory; parallel runs and the topological-sort bug that made scenes start before the scriptwriter feeding them; subtitles built not from my script but from what was actually said on screen; a design system I generated with a single prompt, and packaging rendered in Remotion. Plus an honest look at the numbers — twelve dollars a clip and 22% full-watch retention.
- 07 Video Generator, part 7: testing whether a character stays the same Two characters, a croissant, and experiments in Vini Studio. How I test character consistency, reproduce failures, and revise a hypothesis when the next test contradicts it.
Series · 01–02
CRM на Markdown
- 01 CRM on Markdown, part 1: how to create a CRM in an evening — step by step, with an AI agent that finds leads A step-by-step guide to a working CRM without a database, SaaS or a single dependency: install Claude Cowork or Claude Code, add the md-crm skill in two commands or one ZIP, connect the Chrome extension — and the agent triages LinkedIn search results you scroll, files contact cards and keeps the board honest. You talk to it in plain words: add a contact, log her reply, what's on fire. Same skeleton runs an ATS for hiring. It ends with the final split — what the agent takes over and what stays human — and a way to reach me if you want it wired to your funnel.
- 02 CRM on Markdown, part 2: 27 hours, 23 commits and a catalog of failures The anatomy of a Markdown CRM built with an AI agent on a live project — told through its failures, because that's where every rule came from. Someone else's words recorded as the owner's voice, a ready-made comment posted under a stranger's post, a user's silence explained by the product while the user was moving house. Three breeds of failure — data, process, interpretation — each with its mechanism and price; four decisions I reversed within hours of making them; and the main lesson of the whole exercise: a rule without a guard does not survive, even for the person who wrote the rule two paragraphs above.
Series · 01–11
My Architect
- 01 My Architect, part 1: giving an AI agent memory and a plan An AI agent forgets the project after every session. How My Architect keeps the plan, requirements and decisions alive between sessions: a UI for humans, MCP for agents. First post in a series of seven.
- 02 My Architect, part 2: an architecture with no database Files are the agent-native choice, a database is dead weight here. YAML and Markdown as storage: atomic writes via rename, a per-project mutex after a real lost-update bug, live sync through chokidar and WebSocket.
- 03 My Architect, part 3: a planning model that makes agents think like architects How My Architect models a project so the agent plans like an architect: node hierarchies with presets, four requirement types inherited down the tree, a title linter and a User Story Map the agent reads as text. Third post in the series.
- 04 My Architect, part 4: the agent's working loop How an agent runs a project for weeks: get_next_task picks the work, docs stay current while the code is written, deferred items become nodes, and /reconcile keeps the plan honest against the codebase.
- 05 My Architect, part 5: shipping skills and MCP via the Claude Code marketplace An MCP server, a skill and four slash commands used to take a manual to install. The Claude Code plugin marketplace reduces it to two commands: how the manifest works, why the skill lives in a public repo, and what version discipline costs. Final post of the series.
- 06 My Architect, part 6: the human side — WBS, Building Blocks and the User Story Map The series covered the agent; this part is about me. How an architect uses My Architect by hand: WBS mapped onto building blocks, a User Story Map for prioritization, and why a diagram that an agent executes beats any whiteboard.
- 07 My Architect, part 7: Obsidian, Claude Code docs and OpenClaw — field notes I tried each of them for one reason: to be productive on real projects with AI. What job Obsidian, the Claude Code docs approach and OpenClaw actually do, where each hits its ceiling, and what that taught me about project memory.
- 08 My Architect, part 8: Event Storming as sequence control, not another diagram A picture-diagram checks nothing. Event Storming does: every event has a cause, every command an effect. How an agent runs a board through a deterministic sequence analyzer, closes mechanical gaps to zero, and leaves unresolved business questions visible. On a live Forklift project — 5 contexts, 17 gaps caught.
- 09 My Architect, part 9: Recursive context — teaching the agent to admit what it hasn't read From MIT's Recursive Language Models paper to the recursive-context skill in one day: why a huge context window doesn't solve big data, how a fan-out of isolated sub-agents with honest coverage replaced a Python library, and what only live runs could catch — from a baseline that turned out too good to a 14-millisecond bug.
- 10 My Architect, part 10: a repo memory for the agent — a code graph it's not allowed to trust The recursive-context story continues: four releases in one day teach the agent a persistent code graph (Graphify) under strict distrust rules — freshness gate, facts only from live files. Plus a controlled A/B test of graph vs grep on NestJS: −71% tokens on subsystem understanding, parity elsewhere — and the hidden sub-agent costs that make naive benchmarks lie.
- 11 My Architect, part 11: eight hundred commits I didn't write Not a demo but a live project — a video pipeline with 802 commits in 78 days, with a front end, a back end, video rendering, Python sidecars and five model vendors, and my_architect holding it all together. Inside: the real numbers from the tracker (650 nodes, 296 requirements, 63 RFCs), how a requirement is wired to a commit and a test through a coverage gate that fails the build, why I draw HTML wireframes before any code, what commit timestamps actually measure in the story of “I hand a task over at night and by morning it's done” — the day has two humps with the night between them, and morning commits are twice the size of daytime ones — and why all of this opens a door for teams instead of closing it.