KAIST researchers showed that letting a language-model-driven trading bot rewrite its own prompt every five days lifted its Sharpe ratio from 2.94 to 4.00 and added a 50-basis-point outperformance over a 50-day trial. The gain came from forcing the bot to move every calculation from prose into executable Python code.

Why static prompts fall short

Most LLM agents that execute code start with a fixed system prompt – a block of text that tells the model how to behave. The prompt often treats the code tool as optional decoration. The model may still “think” about volatility or expected returns, but it writes the numbers in plain text and then decides how much to invest based on vague guesses. The result is a disconnect: the model can generate perfectly valid code, yet the final trade size is chosen outside that code, leaving the decision vulnerable to hallucination.

The EvolveTrade approach

The KAIST team replaced the static prompt with a meta-agent that reviews the bot’s performance on a rolling five-day window. After each review the meta-agent rewrites the system prompt, tightening the contract between the language model and its code interpreter. The new prompt mandates three concrete steps:

  • Build a table of metrics for every asset using Python.
  • Apply explicit mathematical formulas to score the assets.
  • Derive target portfolio weights inside the code rather than in markdown text.

In practice the evolved bot called the Python tool 11 times per trading cycle – to compute metrics, validate risk limits, and calibrate weights – whereas the baseline agent called the tool only once, relying on textual heuristics for the rest.

Numbers that matter

During a 50-day back-test on a standard asset universe, the evolved agent posted:

  • Sharpe ratio: 4.00 vs 2.94 for the static agent.
  • Cumulative return: 10.56 % vs 8.88 % for the static agent.

The 50-basis-point outperformance stems directly from the tighter binding of actions to deterministic math. By making the model’s decision pipeline fully observable and repeatable, the researchers eliminated the “guesswork” that typically drags down LLM-driven trading strategies.

Who wins, who loses

For developers building financial bots, the lesson is clear: if the agent has access to a runtime environment, the prompt must compel it to express every actionable insight as code. Ignoring this principle lets the model treat tool outputs as background chatter, which can erode performance and increase risk.

Limits and counter-points

What to watch next

Takeaway: Letting an LLM rewrite its own instruction set forces every trade decision into verifiable code, turning a fuzzy text-based agent into a disciplined, higher-return engine. Developers who ignore the prompt-tool contract risk leaving money on the table.