Skip to content
1 of 3 project slots open
Writing
AI · part 5 of 63 min read

Tools your agent can drive

Why Shelf and Cutscene are built MCP-first: the agent works on a canvas you can watch, and checks its own output.

Meer Habib

Senior Mobile Engineer · Chittagong

Most of my week is spent with an agent in the loop: Claude Code or Codex writing, running and checking code beside me. At some point I noticed that the slowest parts of shipping an app weren't the code any more. They were the things around the code: App Store screenshots, the launch film, the demo video.

So I started building tools for those, and building them for the agent first.

Shelf and Cutscene

Shelf designs App Store screenshots. You give your agent the app's screens; it studies them, writes a design brief, and builds the whole set on a canvas: device frames, type, backgrounds, widgets.

Cutscene makes product demos and launch films. Point it at a screen recording and it frames the footage in a 3D device, finds the moments worth zooming into, and cuts them. Point it at a repo and it writes the film: scenes, motion, type, transitions.

Both are connected to the agent over MCP:

claude mcp add shelf -- npx tsx ./mcp/server.ts

Three rules for agent-driven tools

1. The work happens somewhere you can see it

The agent doesn't generate a file and hand it over. It works on a live canvas in the browser. Every frame, layer and caption appears as it's made. You can watch, and you can take over at any moment.

2. The agent checks its own output

Both tools let the agent look at what it rendered (a screenshot of the finished set, a still from each scene) and fix what it doesn't like before calling the job done. That single loop, render → look → fix, is the difference between a draft and something you'd actually ship.

3. The API is the product

A human-first tool with an API bolted on gives the agent a keyhole. An agent-first tool gives it the same verbs a person has: create a scene, move a layer, swap a font, preview a frame. The UI is then just another client of those verbs, which is also why taking over by hand works so cleanly.

What changes

When the agent can drive the tools, the unit of work changes. It's no longer "make eight screenshots". It's "here's the app, here's what makes it good, tell that story on the store page", and then a review of the result.

I still make every call about what the product is. The agent handles the hundred small decisions after that, and shows me its work.


Both tools are on the bench right now. If you want AI features, or AI tooling, built into your product the same way, start a brief.

Building something like this?

Booking new projects for Q4. Replies within 24h.