Research / Copilot architecture
Technical paper
Tool routing, a live context pipeline and model choice for gold research
A language model on its own cannot know today's gold price. This paper describes how the RM Intelligence copilot connects Claude to live XAU/USD data through the Anthropic API: how tools are chosen, how market context reaches the model, what happens when data fails, and how model size should be matched to the task.
1. The problem: stale knowledge in a live market
A model's training data stops at a cutoff date. Asked "where is gold now?", an unconnected model will either decline or produce a plausible number that could be wrong by hundreds of dollars. In a financial setting a confident wrong number is worse than no number. So the design has one rule: every figure that describes the market now must come from a tool call made during the conversation, carry a timestamp, and be cited.
2. Tool routing: the model chooses, the server executes
We use the tool-use capability of the Anthropic Messages API. Each request sends the model JSON-schema descriptions of three read-only tools. The model decides which ones a question needs. Our server runs them and returns the results, and the model continues from there.
| Tool | Input | Returns |
|---|---|---|
fetch_live_xau_usd | detailed (boolean) | Reference price, observation time, freshness state, source; optional bid/ask and daily change |
get_gold_technicals | timeframe: 1h, 4h or 1d | Last close, SMA20, SMA50, RSI14, ATR14, 20-bar high/low, position versus averages |
get_model_signal | none | Multi-factor regression coefficients with t-statistics, R², residual z-score, model fair value, latest factor levels |
The loop is short and bounded:
for round in 0..4:
response = claude(messages, tools)
if response.stop_reason != "tool_use": return response
for each tool_use block:
result = run_tool(name, input) # server-side, read-only
append tool_result(result) to messages
# final round: tools disabled, the model must answer with what it has
Three design choices matter here:
- Tool descriptions do the routing. Each description says when to call the tool ("before stating any trend, regime, level or volatility figure"). In practice that routes questions more reliably than a separate classifier model, and there is one fewer component to fail.
- Bounded rounds. At most four tool rounds per answer. On the last round tools are switched off, so the model cannot loop forever and has to answer from what it has gathered.
- Read-only by construction. No tool can reach an order-entry or trading account. The research boundary is enforced by what the server can do, not only by the prompt.
3. A retrieval-augmented context pipeline built on live APIs
Retrieval-augmented generation (RAG) usually means pulling passages from a document store into the prompt. For market questions the relevant "documents" are live data feeds, so our retrieval step queries APIs at the moment of the question rather than searching a vector index. The context reaching the model has three layers:
- A static reference pack in the system prompt: role and limits, a numbered source register (S1 to S14), the conditional research playbook and the response policy.
- A workstation snapshot attached as JSON: CFTC positioning, US Treasury real yields and US CPI, each with its own observation time and freshness state.
- Tool results fetched during the conversation: live price, technicals and the model signal.
Every value in every layer carries two timestamps: when it was observed at the source and when we retrieved it. It also carries a state: current, delayed, stale or unavailable. Freshness thresholds depend on the series. A spot quote goes stale in minutes, while a weekly CFTC report is not stale until well past its next scheduled release. The model is told to say when an input is not current.
4. Fail-closed behaviour
When a data source fails, the pipeline says so instead of guessing:
- The price chain tries gold-api.com first and then RM's MT5 broker terminals. If both fail, the tool returns an explicit error ("No value was fabricated") rather than an old cached number presented as current.
- Each input's freshness state is checked against its own threshold, and stale inputs are labelled, never silently reused.
- The response policy makes WAIT or INSUFFICIENT DATA the default when evidence is mixed, stale or incomplete.
We found a real example of why this matters while writing this paper. In September 2026 the CFTC's public data began storing COMEX gold's market code with a trailing space. An exact-match filter stopped finding new weeks, and the positioning input quietly fell back to a July report. Because every value carries its observation date, the daily report flagged it as stale instead of presenting it as current. We now key the query on the CFTC contract code (088691), which is stable.
5. Treating market data as untrusted input
Any text that enters the prompt, whether a user message, a tool result or a data feed, could contain instructions aimed at the model (prompt injection). The copilot's policy treats every user message and tool result as data, never as instructions. Tool outputs are structured JSON built by our server, not raw third-party text. The copilot never asks for credentials, and API keys stay on the server and never reach the browser. Requests are rate-limited per member, and conversation history is capped in length.
6. Model choice: matching model size to the task
Production today: the copilot runs on a single frontier model, Claude Opus 5.5, with an automatic fallback to Claude Opus 5 if the configured model is unavailable to the API key. One capable model handling both tool routing and synthesis keeps the system simple, and each member question is a low-volume, high-value request where answer quality matters more than milliseconds.
Design proposal, not yet deployed: financial AI workloads fall into two classes with different needs.
| Workload | What matters | Suitable model class |
|---|---|---|
| High-frequency, narrow extraction: normalising feed payloads, tagging headlines, checking a quote against a threshold | Latency and cost per call; the output schema is fixed | A small, fast model such as Claude Haiku 4.5 |
| Low-frequency, open-ended synthesis: weighing conflicting macro drivers, writing a conditional research case with invalidation levels | Reasoning depth and calibration; latency matters less | A frontier model such as Claude Opus 5.5 or Claude Fable 5.1 |
A tiered design would give the narrow parsing work to the small model and keep synthesis on the large one. We have not deployed or benchmarked this tiering, so we make no latency or accuracy claims for it. If we adopt it, we will publish measured results here. Tiering also adds a failure point: a small model that mislabels an input can mislead the large one. So any routing layer would need the same observation-time and freshness checks as the rest of the pipeline.
7. How we check it
- Numbers against sources: every price or indicator in an answer should match a tool result from that conversation, and the source tags make this checkable.
- Freshness honesty: when a feed is deliberately made stale, the answer must say so.
- Refusal boundaries: requests for guaranteed outcomes, account-specific lot sizes or order placement must be declined, with a conditional checklist offered instead.
8. Try the pattern yourself
The live-price tool described in section 2 is available as a small MIT-licensed Python package. The developer guide shows how to give Claude a live XAU/USD tool in about twenty lines, and the source is on GitHub.