Agent OS

Next-Gen AI Agents

A UX research experiment exploring how people interact with, delegate to, and build trust in a team of specialized AI agents — and what that taught me about designing for agentic orchestration.

How do people actually use AI agents when given access to a full team of them? That was the question driving Agent OS — a UX research experiment designed to learn how users interact with, delegate to, and develop trust in specialized autonomous agents.

Setting the Scene

Agent OS was a research-driven design project I undertook to understand the emerging space of agentic orchestration. Rather than shipping a polished product, the goal was to prototype a realistic multi-agent platform — six specialized agents with distinct roles, memory, and tool access — and observe how people naturally interacted with it. What tasks would they delegate first? Where would trust break down? What mental models did they bring to the experience?

The Problem

The core research question: if you gave teams access to a full suite of specialized AI agents — not just a single chatbot — how would they use them? Most AI tools at the time were point solutions, siloed and disconnected. I wanted to design the interaction layer for something more orchestrated, and learn what that actually required of the UX.

What made it hard

  • →Designing for agentic systems required rethinking fundamental UX patterns — there's no established playbook for how users should delegate to, monitor, and intervene with autonomous agents
  • →Each of the six agent roles required distinct interaction models — a Research Agent behaves very differently from an Automation Agent, and the UX had to reflect that
  • →Human oversight needed to feel natural, not bolted-on — designing approval gates that didn't break the sense of autonomy was a core tension
  • →Trust calibration was a moving target — users' comfort with delegation shifted rapidly as they observed agent behavior, requiring adaptive interface responses
  • →Understanding when users wanted to stay in control vs. hand off entirely was one of the most interesting and nuanced findings of the experiment

The most important design insight wasn't about features — it was about delegation psychology. Users didn't trust agents because they were capable. They trusted them because they could see exactly what they were doing and why.

What the Research Revealed

  • →Users naturally started with low-stakes, reversible tasks before delegating anything consequential — the UX needed to support a gradual trust-building arc, not assume immediate buy-in
  • →Transparency was the primary trust lever — when users could read an agent's decision log in plain language, their confidence in delegating more grew significantly
  • →Specialized agents created clearer mental models than generalist agents — users understood what a 'Research Agent' did, which made delegation feel less risky
  • →Approval gates weren't friction — they were the feature that made users comfortable handing off higher-stakes work at all
  • →Workflow visibility (seeing what all agents were doing simultaneously) revealed coordination behaviors I hadn't anticipated — users treated the agent dashboard like a team standup

The Approach

I designed Agent OS as a high-fidelity prototype of an agentic operating system — six specialized roles (Research, Design, Sales, Support, Analytics, and Automation), each with its own memory, tool access, and interaction model. The UX explored how a shared Memory Layer, structured Workflow Pipeline, and Human Approval gates could work together as a coherent system. The goal was less about the agents themselves and more about the design of orchestration: how do you give someone meaningful control over a team of autonomous systems?

Six specialized agents, not one generalist

Rather than a single all-purpose agent, Agent OS ships six purpose-built agents with domain-specific training, tools, and prompting. Specialization produced dramatically better outputs and clearer accountability — you always know which agent owns which task.

Four-step workflow pipeline

Define → Connect → Set Boundaries → Deploy & Monitor. Every agent workflow follows the same proven structure. This consistency reduced onboarding time and made workflow creation predictable enough that non-technical users could build production workflows in minutes.

Human Approval as a first-class feature

Every agent workflow includes configurable approval gates — points where a human reviews before the agent proceeds. This wasn't a safety compromise; it was the feature that unlocked enterprise adoption. Control and delegation aren't opposites.

Memory Layer for persistent context

Agents retain context across sessions through a structured Memory Layer. This meant a Research Agent could remember a competitor analysis from last month and reference it intelligently in today's workflow — without re-prompting.

Audit logging as a trust interface

Every agent action is logged immutably with full context: what triggered it, what the agent decided, what it produced, and what happened next. This wasn't compliance checkbox — it was the primary trust-building mechanism for enterprise buyers.

Proving It Out

I ran the experiment with a small cohort of participants across research, operations, and creative roles. I tracked not just task outcomes but behavioral signals — how often did users override an agent? How did that change over time? What tasks did they choose to delegate first vs. last? These patterns informed the core UX findings around agentic trust design.

  • →Users consistently started with information-gathering tasks before delegating anything generative or consequential — the UX needed to scaffold that journey intentionally
  • →Override frequency dropped significantly across sessions as familiarity with agent behavior increased — suggesting that predictability, not just capability, is a key trust driver
  • →The agent dashboard (seeing all six agents' status simultaneously) emerged as an unexpectedly high-value screen — users returned to it frequently to maintain a sense of situational awareness
  • →Participants who understood the Memory Layer — that agents retained context — delegated more complex, multi-session tasks much earlier than those who didn't
6
Specialized agent roles prototyped
3
Core orchestration patterns identified
4-step
Workflow pipeline UX designed and tested
2026
UX research experiment

Beyond the Numbers

  • →The experiment deepened my understanding of agentic orchestration as a design discipline — it's less about UI components and more about trust architecture
  • →Designing the Memory Layer interaction revealed how much users anthropomorphize agents — and how that shapes their expectations of what agents should 'remember'
  • →Approval gates taught me that control and autonomy aren't opposites — the right handoff point is a design decision, not a default
  • →The six-agent mental model gave participants a clearer framework for thinking about AI delegation than any single-agent system I'd previously explored

What I Took Away

This experiment fundamentally shaped how I think about designing for agentic systems. The hard problems aren't technical — they're about legibility, delegation, and trust calibration. Designing for agentic orchestration means designing the human side of the handoff, not just the agent.

  • →Agentic UX is primarily about trust architecture — transparency, predictability, and legible decision-making matter more than feature richness
  • →Specialization creates clearer mental models — users navigate multi-agent systems more confidently when each agent has a defined, well-named role
  • →Approval gates are a design lever, not a safety compromise — they determine where the human stays in the loop and shape how much autonomy feels comfortable
  • →Orchestration visibility (seeing the whole system at once) is a first-class design surface — it should be treated with the same care as any individual agent interface

What's Coming Next

The patterns from this experiment directly inform how I approach agentic product design — particularly around trust scaffolding, delegation UX, and designing systems where humans and agents share a workflow. A current extension of this work is exploring Agent OS as a Markdown skill library — a structured set of .md files that define each agent's role, capabilities, decision logic, and handoff protocols. The goal is to make the system natively consumable by AI agents themselves, so the orchestration layer can be loaded as context rather than hardcoded. It's an exploration of how design documentation and agent instructions can be the same artifact.