Design AI agent architectures that hold up. Even under attack.

Sketch an agent system on a canvas and simulate it: what it costs, where it breaks, and how an attacker gets in. No code, nothing deployed.

The AgentSim Studio running the helpful inbox puzzle: an untrusted email steers the sales assistant, which can send customer records out by email. The finding is ADV-02, the lethal trifecta, critical. Per 1,000 runs 14.16 EUR, p95 10.7 s, under attack it leaks data.
How it works

A demo shows the happy path. AgentSim shows the other four.

Every agent architecture you draw runs against four simulation engines: trace, cost and latency, chaos, and adversarial. You see what a production incident would show you, without the incident.

The trace engine on the refund design: every message, tool call and reply in order, with timings, ending with the answer reaching the customer.

Graded by rules, not opinions. The same design always gets the same result, and every finding names the rule behind it.

Security by design

Find the attack path before someone else does.

Agents read content written by strangers: emails, web pages, uploaded files, customer messages. That is how prompt injection gets in. The adversarial engine traces where that content can travel and flags every route to a tool that leaks data, spends money or changes something for good.

6 attack rules12 Red Team puzzles3 safety patterns
  1. ADV-01Attack pathA model that reads untrusted content can be steered by its author, and so can every tool it calls.
  2. ADV-02Lethal trifectaPrivate data, untrusted content and a way to send data out, all in one model, is how data gets stolen.
  3. ADV-03Poisoned memoryAnything written to memory comes back in later sessions, including an attacker’s instructions.
  4. ADV-04Retrieval ignores permissionsA model can’t keep secrets. Anything retrieved into its context can end up in the answer.
  5. ADV-05Guardrail as the only defenceGuard models catch many attacks, but not all. They lower risk; they don’t remove it.
  6. ADV-06Approval shows a summaryIf the approver sees the model’s description, a hijacked model describes something harmless.
RiskVoid

Already running an AI agent? Secure it with RiskVoid.

AgentSim is where you design an agent and test the design. RiskVoid checks an agent you have already built: it analyzes the source code, points to the vulnerable paths, and shows the code behind each finding and how to fix it.

01Reads your sourceRead-only access to your repository. Agents, prompts, tools and the resources they reach are found from the code itself.
02Maps trust boundariesData flows and trust boundaries are modelled, and every candidate risk path, from injected content to a tool that does damage, is traced.
03Evidence you can fixEach finding links to the lines of code behind it, with analysis coverage and remediation guidance. Rerun after the fix.
For engineers, free

Five ways to get good at agent design

Start with ten-minute lessons, test yourself on open briefs, attack other people's designs, and keep a library of patterns you've actually seen fail.

The Learn screen: every lesson runs the same loop of seeing the problem, building the fix, running it and keeping the takeaway, followed by the Foundations lessons.
AgentSim for Hiring

See how a candidate thinks about agent architecture, with evidence you can explain, in one round.

Coding tests show whether someone can write code with AI. An AI system design interview on AgentSim shows whether they can design an agent system that works, scales, fails gracefully and resists attack.

A live interview: the candidate\u2019s refund design on the shared canvas, with the private interviewer console suggesting a question about a refund that times out after the money moved.
  1. Step 1Pick a templateEight staged problems, from a support agent with refunds to a multi-tenant agent platform, for mid, senior and staff roles.
  2. Step 2Run a live pair-design roundCandidate and interviewer share one canvas, next to your usual video call. You unlock new requirements as the design grows.
  3. Step 3Decide from evidenceA scorecard with engine evidence pre-filled, a replay of the build, and a report the committee can read in five minutes.

Why a design round

Coding testCan they write code with AI?
Whiteboard roundDepends on the interviewer. Thin notes, no shared rubric.
AgentSim roundCan they design an agent system that holds up? With a replay to prove it.

What your team gets

Questions that follow the designA private console turns live findings into follow-up questions, so every interviewer probes the gaps that matter.
One rubric, eight competenciesFrom requirement clarification to safety, with what "meets the bar" looks like at each level.
Replay, not memoryEvery change, run and note is recorded, so the debrief is about what happened.
People decide, alwaysNo automated rejections, rankings or verdicts. EU hosting, a DPA for every customer, retention you control.
Pricing

Learning is free. Hiring teams pay per seat.

LearnFreeFor engineers learning agent design.
  • All lessons, challenges, puzzles and patterns
  • The Studio, with share links
  • No sign-up, nothing to install
Hiring pilotFreeFor teams who want to try a round first.
  • Your first 20 live interviews
  • All 8 templates
  • A short weekly feedback call
Team149 EUR a monthFor a team hiring a few engineers.
  • 2 interviewer seats
  • 15 live interviews a month
  • All 8 templates
Growth399 EUR a monthFor companies hiring steadily.
  • 6 interviewer seats
  • 50 live interviews a month
  • Saved templates, candidate comparison
Questions

Good to know

Do I need to write code?

No. You design with nodes and connections: sources, models, agents, tools and controls. The engines run your design for you. It’s about architecture decisions, not syntax.

Are the costs and latencies real?

They come from a dated calibration table chosen to be in the right range for late 2026. They’re teaching estimates, not quotes from any provider, and the rules behind them are public.

Who is it for?

Engineers building with LLMs who want to design agents that hold up in production, and the teams who hire them.

Does AgentSim decide who gets hired?

Never. The engines produce evidence and interviewers write the recommendation. Nothing ranks, filters or rejects candidates automatically.

Where is candidate data stored?

In the EU. Your company is the controller, we’re the processor, and retention is set by you, 12 months by default.

Will candidates have seen the interview problems before?

Hiring templates come from a private bank that never appears in the free product, and each one has variants that get rotated.

Start with a five-minute lesson. Or run your next design round on AgentSim.

Get early access

AgentSim opens to a first group of engineers soon. Leave your email and we’ll write when your access is ready.

Only about access. No newsletter.