A picture is the wrong input
Give an agent a screenshot and ask for the app, and it will write something. It reads layout well. What it cannot read from pixels is structure: which row opens which screen, whether a list repeats one screen with different content or each row is its own screen, which tab owns a pushed screen, and what the native control is on each platform.
So it guesses, and you review every guess. On a ten-screen app, that review is most of the work.
What the agent receives from the App Builder
When you press Send app to agent, the App Builder checks every screen in the order people reach it, flags anything unfinished, and gives you one line to paste into your agent. The agent reads your design from it and calls emit_app, which hands back:
- every screen in navigation order, with its blocks and their props;
- every arrow as a real route: a tab, a pushed screen or a pop-up;
- the components each screen uses, and the files to write for them;
- an app shell from a skeleton we wrote by hand and tested on real devices on iOS and Android.
If anything is still unfinished, emit_app returns it as a structured flag and refuses to write the app until you fix it or tell it to go ahead. A warning in prose is one an agent can summarise away; a refusal is not.
The shell is solved, not generated
The hard part of a multi-screen app is rarely any one screen. It is the provider order at the root, safe areas, scroll insets under the tab bar, a pushed screen inside the right tab, and the places iOS and Android need different code.
We do not generate that part from scratch for each app. It comes from a small set of skeletons we wrote and verified, and your screens go into one. That is the trade: the shell is ours and fixed, and everything on the screens is yours.
The agent checks its own work
After writing, the agent calls verify_app. It compares the project with what the agent was told to write and reports anything missing or changed, so the agent knows when it is finished instead of announcing success. It also stops the build when your toolchain would make an app that crashes at launch; the compatibility page explains when.
Later, add_screen adds a screen you design afterwards to the right tab, and update_components updates the components and app shell we supplied as iOS and Android move. Your own code is left alone.
You do not need the most expensive model
Everything the agent reads from us is written to be followed step by step, and every tool answer is kept short. We tested a full app build with Claude Sonnet: no top-tier model needed, and our tools used about 2% of the agent’s context. Other vendors’ models have not been through the same test yet, so we do not claim them.
What it does not do
- It does not write your backend. Data, sign-in, payments and push are yours, and the same agent can add them.
- It does not build web apps or tablet layouts. Phone apps for iOS and Android only.
- It does not decide your design. Your agent can draft one for you to edit, but what gets built is what you approved on the canvas.
What you need
You need a Mac with Xcode 26 to build the iOS app. Android needs only the Android SDK. You also need a coding agent, connected once from the Connect your agent page, and a subscription to generate the app: $15 a seat a month, or $12 billed yearly, at the beta price (usually $25 and $20).
Open the App Builder, pick a starter app, and send it to your agent.