Selected Work
AGI (2025) - AI Agent Phones

Android mobile agents are designed to allow an Android agent be able to control and operate a phone. It works by training a computer use model with images of user interface components like dropdowns, radio buttons, cell columns and rows, with text prompts and learns how to operate a computer. The benchmark of computer use agents are at OSWorld. Android agents are a branch off of computer use agents, like Browser Use, Amazon Nova Act, Claude's Computer Use agent, and more. Instead of operating a browser, AGI and these other companies operate the Android OS.
In December 2025, ZTE announced the first Android phone to fanfare and sellout according to the press.
Since AGI wasn't shipping native hardware, we had to build out lots of custom onboarding and permissions for the user to manually activate, before the user had even a chance to then enable the agent. Studying the patterns of Nikita Bier's Explode app, which also had a manual and involved activation process, I designed prototypes to make it as easy as possible for the user to go through the onboarding flow to activate.
In addition to the onboarding, I worked on prototyping the agent navigation interpretability state. The AGI Android agent takes over your phone controls when navigating the device, so I designed a state machine to allow users to steer the agent when it is stuck in a local minima or hallucinating.
Ultimately, I believe we are still years away from AI Agent phones truly entering the mainstream. As a consumer, I expect 99.9999% uptime on Instagram, Gmail, Facetime, Zelle, Robinhood, or any consumer app with traction. Android agents are the same. When testing AGI internally, I saw accuracy rates of less than 99.999%.
With the expectations of consumers in the world today of <200 ms for loading, infinite scrolling, streaming text generation from LLMs, and 99.999% uptime for services like Github, Gmail, Instagram, I believe Android phone agents are still on a long horizon out, even if it meets the standards of commercial consumer software: 99.999% reliability and <200 ms response time.
The user knows what it wants better than the android agent operating on behalf of a prompt that was ill formed. According to Claude Shannon's model of information transmission, the transmitter submits a message, however before the receiver can receive the signal, there is a noise source that interferes with the signal. In the analogy of AI Android agents, the user is the transmitter and the Android Agent is the receiver. The prompt box is the noise source that captures the signal but also adds noise.
Because LLMs hallucinate and are not pure functional machine learning algorithms like supervised or unsupervised learning algorithms, the noise overwhelms the signal. Compare that to the case when the transmitter is the user and the Android phone is the receiver. The noise source now is only the hand and fingers interacting with the phone. The noise input are the errors that arise when mistapping, or gaps in mental models of the user interface and information architecture. The surface area of these tactile noises are exponentially smaller compared to the noise input of Android Agents.
Agent States

Gradient glows for agent progress

Internal tracing dashboard (2025)

Krea (2023)
In 2023, Midjourney had just launched on Discord and become the biggest server on Discord, with hundreds of millions of users eager to try the latest text-to-image model. Its status quo was using slash commands to generate images, with the image generation command being "/imagine".
Due to the success of Midjourney, a whole host of other text-to-image companies sprung up. There was Leonardo.ai, getimg.ai, Playground AI, and even Adobe Firefly. Krea's take at that time was to use the canvas as a differentiator to Midjourney and allow users to not just generate images, but then visually organize them on a web canvas with further editing and refinement.
Also in February 2023, the paper, "Adding Conditional Control to Text-to-Image Diffusion Models" aka Controlnet came out. This craze led the stable diffusion community and text-to-image companies to race toward building out workflows supporting this new model.
Inpainting and outpainting are simple concepts that let people draw on a specific part of an image and then use text prompts to generatively augment the picture. They already existed before the launch of Controlnet, but Controlnet allowed users to guide generation with edges, depth maps, pose skeletons and more.
I designed different information architectures for interacting with images on the canvas + prompting. Along with entrypoints for prompt + image + mask and visualizing loading and error states.
The space in 2026:
The space has bifurcated into 2 workstreams and companies marketing themselves to each:
- The generic consumer workstream that is dominated by Midjourney, Higgsfield AI, Capcut and the large labs - Google Gemini, ChatGPT, Claude.
- The prosumer workstream that is dominated by a ComfyUI node-based interface. This space has players like Weavy (acquired by Figma), Flora, ComfyUI, Krea's node editor.
In 2023, the market had the unbounded belief that everyone was going to be an AI artist and the newness of trying text-to-image generation felt incredibly novel at the time. Now with the fairy dust worn off, it is incredibly clear that not everyone will be text-to-imaging on a daily cadence or even weekly cadence, and the TAM of the text-to-image space is set. Midjourney itself knows this, milked the cash cow while it was on top of the craze and has since moved onto medical spas in San Francisco.
Krea had $8M ARR as of April 2025, and with a valuation of $500M that leads to a 62.5 revenue multiple. Comparing this to the other software acquisitions of 2026 in this space:
- Weavy getting acquired by Figma for rumored ~$200 million with only a valuation of $13 million.
Krea raised $47 million on a post-money valuation of $500 million, ComfyUI raised $30 million also on a post-money valuation of $500 million, Flora raised $42 million on a probably similar $500 million valuation.
My hypothesis is that similar to how Clay pioneered the GTM engineer, Palantir pioneered the forward-deployed engineer, and if Krea, Flora, and ComfyUI can pioneer a new role like "forward deployed creative" and build up the recruiting motion to place these into companies, it can build a new lever for enterprises to spend money to hire these marketing / creative magicians.
Exciting times to be in AI creativity.

Boom times in San Francisco - The Information (Press)
Boom times in San Francisco - The Information (Press)

