Neko.AI: designing trust for an asynchronous code review agent
When an agent works in the background, clarity of state is the product. Without it, capability reads as failure.
Business context
Neko.AI helps engineering teams review code by connecting a repository or cloud environment and letting an agent inspect it. The founding team had proven the core capability. What was missing was an experience strategy that made that capability understandable, credible, and recoverable when work was slow, queued, or failed.
I joined as sole lead product designer, reporting to the CEO. Audiences ranged from individual engineers to managers and multi seat organisations. Each group needed the same underlying promise: I always know what the agent is doing, and I know what to do next.
Strategic friction
Agentic products break the request and response mental model. Review jobs are asynchronous. Outcomes are probabilistic. Errors are often environmental (permissions, connectivity, service health) rather than form validation. In that setting, interface silence is not neutral. Users invent a story, and the story is usually that the product is broken.
Across the connect and review journey, the experience lacked a durable feedback layer, a consistent information hierarchy for primary actions, and designed responses for empty, waiting, and degraded conditions. Enterprise entry paths also lagged behind who the product was trying to win.
Commercially, that showed up as stalled evaluation, weak retention, and demos that could not carry trust. The model could perform. The product could not communicate performance. That gap is an experience strategy problem before it is a visual design problem.
How I investigated
I ran a structured product audit with the founders: walk the live flows as each audience would, map every moment the agent leaves the user’s sight, and stress demos the way a sceptical technical buyer would. The happy path was credible. The uncertain path was undefined.
The organising insight was simple and consequential. People were not rejecting artificial intelligence. They were rejecting opacity. Trust in an agent is built from continuous system communication. Status, confirmation, severity, and recovery are how users form a mental model of what the agent is, what it is doing, and whether it is safe to depend on.
Insight to opportunity
- Opaque outcomes pointed to a shared feedback language for success and failure
- Invisible background work pointed to an explicit job lifecycle in the interface
- Undefined empty and degraded moments pointed to a state system, not one off screens
- Competing actions and weak hierarchy pointed to a design system with usage principles
- Limited organisational entry pointed to an authentication story that includes enterprise access
Experience strategy
I reframed the work from “improve screens” to “define the experience contract for an asynchronous agent.” That contract answers one question at every step: what is the system doing, and what should the user do next?
My ownership. End to end responsibility for visual language, components, interaction states, layout hierarchy, and governance so a small founding team could ship without recreating inconsistency. Priorities were set with the CEO against demo risk and early retention.
State before spectacle
Waiting, running, success, failure, empty, and degraded are primary product surfaces. They are designed, named, and reused before net new feature chrome.
System over one offs
Tokens, components, and patterns exist so velocity does not fragment the experience. Custom components are introduced only when the agent’s behaviour needs a new interaction grammar.
Governance is part of design
Usage rules for hierarchy, severity, and primary actions protect quality when production is fast and AI assisted tooling is in the mix.
Design recovery as carefully as success
Evaluation and demos fail on the ugly path. Recovery, empty, and outage moments are core experience, not backlog leftovers.
Phasing decision
I prioritised the review loop’s trust layer first (feedback, progress, empty and failure), then expanded system coverage. Spreading effort evenly would have left the highest risk moment unchanged: a buyer still unable to tell whether the agent was alive.
01. Make agent lifecycle legible
Transient, low contrast failure cues told me the interface never owned the job lifecycle. I redesigned confirmation and error as persistent, severity aware communication, then defined progress for queued and running reviews so absence of motion no longer read as a crash.
Decision. For AI products, status is a first order experience requirement. A feature is incomplete until a user can answer, in one glance, whether work is waiting, running, complete, or failed. Model throughput without that answer is not a shipped product experience.
02. Build a system that scales judgment
I rebuilt the visual foundation (palette and tokens), established a component library, and corrected layout hierarchy so primary paths stop competing with secondary noise. Where agent behaviour needed patterns a generic kit cannot express, I designed purpose built components.
Decision. A library without principles is inventory. I paired components with usage governance: what reads as primary, how severity is expressed, which patterns belong on which surfaces. That is how strategic UX protects coherence when a two person founding team ships quickly, including with AI assisted production.
03. Design the paths evaluation actually hits
Buyer journeys fail on moments teams rarely design first: empty connections, permission issues, long running queues, partial results, and service interruption. I defined a state gallery for those conditions so the product answers with intent instead of a blank canvas.
Decision. In AI UX, recovery design is a growth surface. If a team cannot show what happens when the agent stalls or the service fails, they cannot sell reliability. Designing those states turned “what if it goes wrong?” from improvisation into product.
Current outcome
The product is still in build. Metrics and founder feedback are being collected. What already changed is the experience foundation required for credible evaluation: a feedback model, a shippable system with governance, and designed coverage for asynchronous and failure prone moments.
Next, I am aligning evidence to the strategy: qualitative clarity in demos, conversion through first successful review, reduction in “is it working?” confusion, and time to correctly interpret job state.
Evidence I am capturing
- Founder and buyer clarity in live walkthroughs
- Completion from connection through first review
- Volume of status ambiguity questions
- Time for a new user to correctly read job state
Reflection
AI tooling can accelerate interface production. It cannot substitute experience strategy: hierarchy, trust signalling, and recovery under uncertainty. Shipping agent capability without that layer creates products that perform technically while failing commercially, because users cannot form a reliable mental model.
My lead approach on AI UX is to make system behaviour legible, encode craft into a governed system a small team can keep coherent, and design for the moments when the agent is slow, uncertain, or wrong. That is the difference between a capability demo and a product people can evaluate with confidence.
Strategic UX for AI is less about generating screens and more about designing how trust survives uncertainty.