Market manipulation by machines · Part 3

When the evidence is a simulation

Parts 1 and 2 argued from principle. A machine can reach the manipulative pattern without the intent the law requires. Its reason cannot be read after the fact. This part sets out the studies behind those claims, in pictures. They are experiments and simulations, not enforcement cases. That gap is part of the argument. What they show is concrete: how a reward produces a behavior, and how autonomous agents have been made to collude with no word passing between them.

Part 3 of a three-part series on market manipulation and collusion by machine actors. Part 1: When the rogue actor is itself a model. Part 2: When the reason cannot be read.

August 9, 2026

Part 1 and Part 2 made an argument about what the law can and cannot do when a trading model manipulates a market or slips into collusion. The argument rests on research. That research is worth seeing directly. This part draws on four studies. One shows how the choice of reward shapes what an agent does. The other three show autonomous agents reaching collusion with no agreement between them. None is an enforcement case. All are experiments. That is the first thing to say about them.

How the reward makes the trader

Before an agent can manipulate or collude, someone has to build it. The choice that shapes it most is the reward, the measure of success it is set to make as large as possible. A market-making study makes this concrete. An agent rewarded on raw profit learned to hold a large one-sided position and bet on the price. That is speculation, not market-making. Its training was also unstable. The fix was to change the reward. Keep every loss at full weight. Discount only the profit that comes from holding inventory. The agent then kept its position small and neutral. It did the job it was built for.1 The point is general: the behaviors Parts 1 and 2 describe begin in the reward.

Figure 1 · reward design in a market-making agent

The dampened profit-and-loss reward and its effect A losing step is counted in full; a step that profits from holding inventory is counted less, shaved by a fraction eta, a dial from zero (shave nothing, same as raw profit) to one (shave off all holding profit). Under raw profit the agent hoards a big one-sided inventory and is unstable; under the dampened reward it keeps a small neutral inventory and is stable. THE REWARD RULE reward the result of each step, but not equally A step that loses money kept in full A step that profits from holding inventory shaved off counted less DP = ΔPnL − max(0, η · ΔPnL) Take the step's profit or loss. If it is a profit, shave off a fraction of it (that fraction is η). If it is a loss, keep all of it. η (say "eta") is just a dial: how much of the holding profit to shave off. a little is enough 0 = shave nothing (raw profit) 1 = shave off all holding profit The reward you choose changes what the agent learns to do: IF REWARD = PROFIT ONLY big one-sided pile Hoards a large one-sided position. Bets on price. Learning is unstable. IF REWARD = DAMPENED PROFIT small, flat inventory Keeps inventory small and neutral. Earns the spread. Learning is stable.
How the reward shapes the agent. The reward keeps every loss at full weight but shaves a fraction off gains that come from holding inventory. Sitting on a large position then stops paying enough to be worth the risk. The agent shifts from hoarding a big one-sided position to keeping inventory small and neutral. Its learning becomes more stable. The reward rule is from Spooner, Fearnley, Savani, and Koukorinis, "Market Making via Reinforcement Learning" (2018). The inventory sketches are schematic.

How collusion holds without a pact

The studies of machine collusion share a mechanism, not a conversation. The agents settle into a discipline of mutual restraint. Each holds back, keeping the price above the level competition would set. If a rival breaks ranks and the price crosses a threshold, the others revert to hard competition for a finite punishment phase. Cooperation then resumes. The threat of that punishment holds the restraint in place. This was shown first for pricing algorithms,2 then for reinforcement-learning traders in a model market. They sustained the pattern with no agreement, communication, or intent.3 No message is sent, and no deal is struck.

Figure 2 · the price-trigger discipline · price over time

The price-trigger discipline that sustains collusion without a pact Price over time. Restraint holds the price at a high collusive level. A rival breaking ranks triggers a drop to the competitive level for a finite punishment phase, after which cooperation resumes. The threat of the punishment phase, not any agreement, keeps the price high. PRICE TIME → COLLUSIVE LEVEL COMPETITIVE LEVEL a rival breaks ranks restraint holds the price up PUNISHMENT PHASE (FINITE) cooperation resumes No agreement and no message pass between the agents. The threat of the punishment phase is what keeps the price high.
Collusion without a conversation. Each agent restrains its own trading to hold the price above the competitive level. A rival that breaks ranks and moves the price past a threshold triggers a finite return to hard competition, the punishment phase. Cooperation then resumes. The discipline is held in place by the threat of that phase, not by any agreement, message, or shared instruction. Schematic. It shows the shape of the mechanism, not data from one study.

Three demonstrations

The same pattern turns up across very different setups. It is not confined to bespoke trading algorithms. Agents built on a consumer large language model reach the same collusive prices. The degree shifts with small changes in the wording of their instructions.4

Table 1 · three studies, one pattern

StudySetupWhat emerged
Calvano and colleagues, 2020 Repeated price competition in a standard economic model, run by Q-learning pricing algorithms Prices held above the competitive level, sustained by a finite punishment phase, with the agents never instructed to collude
Dou, Goldstein, and Ji, 2025 A model market with the informed speculators replaced by reinforcement-learning traders Collusion sustained without agreement, communication, or intent
Fish, Gonczarowski, and Shorrer, 2024 Repeated pricing by agents built on a consumer large language model The same supracompetitive prices, with the degree shifting on the wording of the instructions
One pattern, three settings. Different markets, different kinds of agent, the same outcome: prices held above the level competition would set, reached by trial and error rather than by instruction. Q-learning, in the first row, names a method by which an agent learns from experience which action pays best over the long run. Supracompetitive means above the price a competitive market would produce.

What the studies do not show

Two limits belong on the record. First, these are controlled experiments and simulations, not real markets under real conditions. Economists remain divided over how robust the collusion is once markets are noisy and crowded. Second, no jurisdiction has yet sanctioned purely autonomous collusion. The behavior has been shown to be possible. It has not been caught, charged, and proven in a live case. That is the honest state of the evidence. It is also why the argument of Parts 1 and 2 matters. The doctrines are being tested against conduct the studies say is coming, before a case arrives to force the question.

Where this leaves the series

Put the three parts together. A machine can reach the manipulative or collusive pattern with none of the intent the law is built to find. Its reason for doing so cannot be read out of the model after the fact. The studies show the behavior is not hypothetical. That is why the most concrete response in Part 2 did not try to read the machine at all; it governed the action instead. Each proposed trade is checked against an explicit mandate before it executes. The record is kept at the boundary. The reason stays inside the box. The account, as this series has argued throughout, has to be built outside it.

Notes

Links captured and verified August 2026. The studies below are experimental and simulation work. No jurisdiction has yet sanctioned purely autonomous tacit collusion. The robustness of the effect in real markets remains contested.

  1. On a market-making agent whose behavior turns on the reward, where a raw profit-and-loss reward induces speculative inventory-holding and an asymmetrically dampened reward restores neutral market-making, see Thomas Spooner, John Fearnley, Rahul Savani, and Andreas Koukorinis, "Market Making via Reinforcement Learning," Proceedings of the 17th International Conference on Autonomous Agents and Multiagent Systems (2018): 434 to 442; arXiv:1804.04216: arxiv.org.
  2. On Q-learning pricing algorithms that learn supracompetitive prices, sustained by a finite punishment phase and a gradual return to cooperation, without being designed or instructed to collude, see Emilio Calvano, Giacomo Calzolari, Vincenzo Denicolò, and Sergio Pastorello, "Artificial Intelligence, Algorithmic Pricing, and Collusion," American Economic Review 110, no. 10 (2020): 3267 to 3297: aeaweb.org.
  3. On reinforcement-learning traders that sustain collusion without agreement, communication, or intent, see Winston Wei Dou, Itay Goldstein, and Yan Ji, "AI-Powered Trading, Algorithmic Collusion, and Price Efficiency," National Bureau of Economic Research Working Paper 34054 (2025): nber.org.
  4. On pricing agents built on a consumer large language model reaching supracompetitive prices, with the degree shifting on the wording of their instructions, see Sara Fish, Yannai A. Gonczarowski, and Ran I. Shorrer, "Algorithmic Collusion by Large Language Models," arXiv:2404.00806 (2024): arxiv.org.

The research described here is current to August 2026 and continues to change as the experimental literature grows and as regulators and market-abuse authorities take up the question of autonomous agents.

Responses from readers

This website does not host open comments. Verified responses are published here at the editor's discretion. Submit a response to editorial@tradesreconstructed.com.