How to measure an embedded AI agent's impact on an exchange

The way to measure an embedded AI agent is a controlled pilot. Randomly assign eligible traders to an agent group or a holdout, then compare them. The primary metrics are funded-to-first-trade conversion, second-trade rate, trading retention and feature adoption. Guardrails check that nothing goes wrong along the way: advice-boundary pass rate, order confirmation rate, and order error and cancel rates.
This guide is for exchange and broker product, growth and data teams. It covers pilot design, the metrics, and a sample timeline. It contains no results, because results depend on the venue.
Why a holdout, not a before-and-after
Trading activity tracks the market. A volatile month lifts trading for everyone. A quiet one lowers it. A before-and-after comparison cannot separate the agent’s effect from that background.
A randomly assigned holdout solves this. Both groups face the same market, the same campaigns and the same app releases. The only systematic difference is access to the agent.
Pilot design, step by step
- Define eligibility. Decide who enters the pilot: all traders on a platform, new funded accounts only, or a set of countries. Exclude groups where the agent is not yet available.
- Randomise at the user level. Assign each eligible trader once, at first eligibility, and keep the assignment fixed. Randomising by session lets the same trader see both experiences.
- Choose the split. A 50/50 split gives the most statistical power. Teams that want to limit exposure early sometimes start at 10% agent, 90% holdout, then widen.
- Pre-register the metrics. Write down the primary metric, the secondary metrics, the guardrails and the decision rules before the pilot starts.
- Analyse by assignment. Compare everyone assigned to the agent with everyone in the holdout, whether or not they opened the panel. This intent-to-treat view avoids bias from comparing only enthusiastic users. Usage-based cuts can be added as secondary analysis.
- Check the plumbing first. Before the real read, confirm both groups look alike on pre-pilot behaviour, and that events fire correctly for both.
Sizing the pilot
The smaller the effect you want to detect, the more traders you need. A rough rule of thumb for comparing two conversion rates, at 80% power and 5% significance, is:
Traders per group ≈ 16 × p × (1 − p) ÷ d²
Here p is the baseline rate and d is the smallest absolute difference worth detecting.
Arithmetic example, illustrative numbers only. Baseline second-trade rate p = 0.30. Smallest difference worth detecting d = 0.03, meaning 30% to 33%.
16 × 0.30 × 0.70 ÷ 0.03² = 16 × 0.21 ÷ 0.0009 = 3.36 ÷ 0.0009 ≈ 3,733 traders per group, or about 7,467 in total.
Use the venue’s own baseline. A data team will refine this with a proper power calculation.
Primary metrics
Pick one as the headline and treat the rest as secondary.
| Metric | Who it applies to | Definition |
|---|---|---|
| Funded-to-first-trade conversion | New funded traders | Share who place a first trade within a set window after funding |
| Time to first trade | New funded traders | Median time from funding to first trade |
| Second-trade rate | First-time traders | Share making a second trade within a set window |
| D7 / D30 trading retention | All enrolled traders | Share trading in the day-7 or day-30 window |
| Feature adoption | All enrolled traders | Share using a feature for the first time: limit orders, alerts, perpetuals |
For definitions and windows, see exchange trader retention and the second trade and why new traders stall before their first trade.
Guardrail metrics
Guardrails do not show that the agent helps. They show that it does no harm while being measured. Each needs a threshold, set in advance, that pauses the pilot if crossed.
| Guardrail | What it checks | How to measure |
|---|---|---|
| Advice-boundary pass rate | The agent stays on the information side of the line | Share of sampled conversations passing human review, plus automated checks on every answer |
| Confirmation on executed orders | Nothing executes without the trader’s confirm | Share of agent-drafted orders that executed with an explicit confirmation; the target is all of them |
| Draft-to-confirm rate | How often drafts match intent | Share of drafted orders the trader confirms; a sharp drop suggests drafts are wrong |
| Order rejection rate | Drafts are valid for the venue | Share of confirmed agent orders rejected by the order API |
| Quick-cancel rate | Orders were what the trader meant | Share of agent orders cancelled or reversed within a short window |
| Support contacts | No new confusion | Tickets mentioning agent orders, per 1,000 enrolled traders |
The draft-to-confirm rate is read both ways. Very low suggests drafts miss intent. It is not a target to push up, because the trader declining a draft is the system working as intended. See why confirm-by-default matters.
For how the advice boundary is tested before launch, read how AI trading agents avoid investment advice.
A sample timeline
This assumes enough traffic to reach the sample size in six weeks. Adjust to the venue.
| Week | Activity |
|---|---|
| 1–3 | Integration, event instrumentation, metric definitions written and agreed |
| 4 | Small ramp; check events fire for both groups and groups look alike; guardrails watched daily |
| 5–10 | Full enrolment at the chosen split |
| 7 | Interim guardrail review; primary metrics not read early |
| 11 | Read conversion, time to first trade and 7-day second-trade rate for all cohorts |
| 15 | Read D30 retention, once the last cohort enrolled in week 10 reaches day 30 |
The D30 read comes later than the others because week 10’s cohort needs 30 more days, just over four weeks.
Common mistakes
- Reading early and often. Checking the primary metric daily and stopping at the first good day inflates false positives.
- Comparing users with non-users. Traders who open the agent differ from those who do not. Compare by assignment.
- Changing the agent mid-pilot. A large change resets what is being measured. Log every release.
- Leaking the holdout. Marketing that shows the agent to holdout traders blurs the comparison.
- Dropping the guardrails after launch. Advice-boundary and confirmation checks belong in production, not only in the pilot.
How Hippo pilots work
Hippo pilots are controlled and measured, for example with holdout groups. Integration takes weeks, not quarters. Every Hippo order is drafted for the trader to confirm on Hippo’s order ticket. Nothing executes until they do, so the confirmation guardrail is built into the product. Passing Hippo’s advice-bait test suite, run against more than one million questions, is a launch gate.
For build-or-buy questions, read build vs buy an AI trading agent and the AI trading agent vendor checklist. To discuss a pilot, see askthehippo.com.
Frequently asked questions
Why does a pilot need a holdout group?
Trading activity moves with the market. Without a holdout, a rise in trading during the pilot could come from volatility, a campaign or seasonality. A randomly assigned holdout experiences the same conditions without the agent, so the difference between the groups isolates the agent's effect.
What metrics show whether an AI trading agent is working?
For new traders, funded-to-first-trade conversion and time to first trade. For all traders, second-trade rate, D7 and D30 trading retention, and adoption of features such as limit orders, alerts or derivatives. Each is compared between the agent group and the holdout.
What are guardrail metrics for an AI trading agent?
Guardrails check that the agent does no harm while it is measured. They include the advice-boundary pass rate on reviewed conversations, the share of executed agent orders with explicit confirmation, order rejection and quick-cancel rates, and support contacts about agent orders.
How long does an AI agent pilot take?
It depends on traffic and the metrics chosen. Conversion and 7-day metrics can be read within weeks of enrolment ending. A D30 retention read needs the last enrolled cohort to reach day 30, which adds about a month.
Hippo provides information, not investment advice.
Part of our guide: Why new traders stall before their first trade, and how to measure it