Skip to content

How to evaluate an AI trading agent vendor: 20 questions

Team Hippo · PUBLISHED · 4 MIN READ

To evaluate an AI trading agent vendor, ask how the agent stays on the information side of the advice line, whether order confirmation is mandatory, how answers cite live data, and what integration needs from your engineers. Then ask how trader data is handled and how a pilot is measured. In every case, ask for evidence rather than descriptions.

The 20 questions below are grouped by area. They work in an RFP, a first call, or a trial. A strong vendor will answer each one specifically. Vague answers are themselves useful information.

Safety and the advice boundary

The advice boundary is the highest-risk area. Background is in how AI trading agents avoid investment advice.

  1. What is the agent forbidden from saying, and how is that enforced at output time? Look for a concrete list of blocked patterns plus output checks, not only a system prompt.
  2. How large is your advice-bait test suite, and what does it cover? Ask about direct requests, leading questions, role-play, multi-turn pressure, slang and every asset class you list.
  3. Is passing the suite a release gate? A report produced after release is not the same as a gate that blocks one.
  4. How does the agent answer “should I buy this?” Ask for live examples. A good answer declines the verdict and still gives the facts.
  5. How are boundary failures found in production handled? Look for a loop from production conversations back into the test suite.

Confirmation

  1. Can the agent ever submit an order without explicit trader confirmation? The answer to look for is no, with no setting to change it.
  2. Which fields does the order ticket show before confirmation? Compare against the list in what an AI-drafted order ticket must show.
  3. How does the agent resolve relative quantities like “half” or “the rest”? Ask to see the quantity echoed in the trader’s own terms.
  4. Where does confirmation happen? On the agent’s own ticket, in a native form, or in another app. Each extra hand-off is a chance for details to drift.

Data provenance

  1. Where does each answer’s data come from? Your venue’s live data, third-party feeds, or the model’s training data. Ask how each is labelled.
  2. Does every answer show its sources and a timestamp? Market answers without a time are hard to trust and hard to audit.
  3. What does the agent do when data is missing or stale? Look for an explicit “I don’t have current data for that” rather than a guess.

Integration

Patterns are compared in thin-client SDK vs deep integration.

  1. What does integration require from our engineers? Ask for the list: market data, account and order APIs, authentication, product metadata, a mount point.
  2. Does integration change our native screens? A self-contained panel leaves them as they are; deep integration does not.
  3. How long does a typical integration take, and what usually slows it? Ask for the dependencies on your side, not just the vendor’s estimate.

Security and data handling

  1. How is the agent’s session tied to the logged-in trader? The agent sees only that trader’s data, and only while the session is valid.
  2. What trader data is stored, where, for how long, and who can access it? Ask for the retention policy and the deletion process in writing.
  3. Is trader conversation data used to train models, and can that be switched off? Get the answer in the contract, not only on a call.

Commercial and pilot

  1. How is a pilot structured and measured? Look for a controlled design with a holdout group and agreed metrics. See measuring AI agent impact at an exchange.
  2. Whose side is the agent on? Confirm it is loyal to your venue: your accounts, your order API, and no routing of traders elsewhere.

How to use the answers

A simple scoring approach works well.

  1. Mark each answer as evidenced, described, or missing. Evidenced means you saw it working or received documentation.
  2. Treat questions 1 to 3, 6 and 20 as gating. A missing answer on any of them is a stop, whatever the score elsewhere.
  3. Run your own bait questions in a trial environment. Write twenty that reflect how your traders actually talk.
  4. Bring compliance in early. Vendor testing does not replace your own legal review in each market.

Answers that deserve a follow-up

Some responses come up often in evaluations. None is disqualifying on its own, but each needs a second question.

Answer you hear Follow-up to ask
“The model is instructed not to give advice.” How is that tested, at what scale, and does it gate releases?
“Confirmation can be configured.” Can it be switched off? By whom?
“Answers use real-time data.” Real-time from which source, and is the timestamp shown on every answer?
“Integration is simple.” Which of our APIs, which auth method, and what changes in our app?
“We take security seriously.” What is stored, for how long, and is it used for training?
“Pilots usually show strong results.” Against what holdout, on which metrics, over what period?

The pattern is the same each time. A description of intent is a starting point. The follow-up asks for the mechanism, the evidence, or the contract term behind it.

If you are still deciding whether to buy at all, start with build vs buy: AI trading agents.

For background on the category, read what an embedded conversational trading agent is and why confirm-by-default matters. More in embedded agents.

Hippo is built for exchanges and brokers evaluating exactly these questions; see askthehippo.com.

Frequently asked questions

What is the most important question to ask an AI trading agent vendor?

Whether the agent can ever submit an order without the trader's explicit confirmation. The answer to look for is no, with no setting that turns confirmation off.

How do you check a vendor's advice-boundary claims?

Ask for the size and coverage of their advice-bait test suite, the current pass rate, whether it gates releases, and a sample of failures and how they were fixed. Then run your own bait questions in a trial.

What should a pilot of an AI trading agent measure?

Outcomes that matter to the venue, such as activation to first trade, repeat trading and answer accuracy, compared against a holdout group that does not see the agent, over an agreed period.

Does a vendor's testing make the agent compliant in my market?

No vendor test replaces your own legal review. Testing shows the agent is designed to stay on the information side of the line; each market still needs its own assessment.

Hippo provides information, not investment advice.

NEWSLETTER

Conversational trading, in your inbox

New guides and research from Team Hippo. No spam; unsubscribe from any email.

By subscribing you agree to receive emails from Ask the Hippo.