Blog

Why Sonnet 4.6 kept dropping a simple format rule

13 min read

I asked five models to do a small agent task while following one rule: start every response with STATUS: 1, then STATUS: 2, then STATUS: 3, and keep counting until the conversation ends.

I expected this to be the warm-up. It seemed almost too simple given what these models can do, but I wanted to see whether I could trip them up systematically with a modest testing budget.

Four models followed the rule throughout every run I tested. Sonnet 4.6 sometimes dropped it within the first few turns, then carried on with the rest of the task as though the rule had never existed. At 200 sequential tool calls, that happened in about 80% of runs.

The five-model comparison covered 230 trajectories and cost about $40 in API credits. The most useful finding came from reading individual runs with halftrace (opens in a new tab): the average scores made this look like a different problem from the one I actually had.

What I measured

The task was find_and_synthesise: make a sequence of tool calls, remember some facts along the way, then submit an answer. I varied the trajectory length from 5 to 200 sequential tool calls and checked four kinds of behaviour.

The instruction_decay probe checked the STATUS: rule. The other probes checked whether the agent remembered earlier facts (state_amnesia), avoided calling tools again with identical arguments (tool_repetition), and issued real tool calls instead of describing them in text (narration_substitution).

Compliance scores run from 0 to 1. A score of 1.0 means the rule was followed on every assistant turn; 0.0 means it was never followed. I'll use a few terms throughout:

  • N is the number of sequential tool calls in a run.
  • A cell is one model at one value of N, such as Sonnet at N=100.
  • A rep is a complete run through that cell. I ran 10 reps per Sonnet cell. For the other models, I ran 10 at the headline cells and 5 elsewhere.
  • A per-cell mean averages the scores across those reps.
  • The Perfect column counts reps scoring at least 0.95. Where the mean is 1.000, every scored turn complied.

For example, Sonnet's score of 0.16 at N=200 means that, across ten separate conversations of that length, an average of 16% of assistant turns began with the correct prefix. It doesn't tell us when the failures happened.

That distinction matters. An average of 0.5 could describe five perfect runs and five abandonments, ten partly compliant runs, or gradual deterioration in each run. I'd investigate those very differently.

One model kept dropping the rule

Each cell below shows compliance at that model's deepest tested N. GPT-4.1 and GPT-4o were tested to N=100; the longer tests reached N=200.

Modelstate_amnesiainstruction_decaytool_repetitionnarration_substitution
Opus 4.71.001.001.001.00
Sonnet 4.61.000.16 (at N=200)1.001.00
Haiku 4.51.001.001.001.00
GPT-4.11.001.00 (to N=100)1.001.00
GPT-4o1.001.00 (to N=100)1.001.00

Nineteen of the twenty cells scored 1.00. Opus 4.7, Haiku 4.5, GPT-4.1 and GPT-4o made no mistakes on these probes at any tested length, across 150 combined trajectories. Haiku, despite being the smaller model, was more reliable than Sonnet on this particular task.

I had pushed the longer runs towards 200 tool calls looking for gradual decay. The other four models gave me very little to investigate.

Sonnet's means, with ten reps per cell, looked like this:

N=5    0.93   ┃██████████████████░░  early near-perfect
N=10   0.58   ┃███████████░░░░░░░░░  ← early uncertainty dip
N=25   1.00   ┃████████████████████  locked in
N=35   1.00   ┃████████████████████  locked in
N=50   0.90   ┃██████████████████░░  lower mean
N=70   0.81   ┃████████████████░░░░
N=100  0.51   ┃██████████░░░░░░░░░░  40% abandon
N=200  0.16   ┃███░░░░░░░░░░░░░░░░░  ← 80% abandon

The dip at N=10 still bothers me. Sonnet did worse on ten-tool conversations than on twenty-five-tool ones. When I first plotted it, I assumed I'd got the scoring wrong. I checked it several times and kept getting the same result.

The decline from N=50 is easy to misread. Each point comes from a separate set of runs at that length. A lower mean in the longer runs doesn't tell us whether compliance faded during any individual conversation.

Here are the individual scores at the deeper cells:

NPer-rep instruction_decay scoresabandon rate
501.00, 1.00, 0.04, 1.00, 1.00, 1.00, 1.00, 1.00, 1.00, 1.0010%
1000.51, 0.02, 0.02, 0.02, 0.03, 0.51, 1.00, 0.99, 1.00, 0.9940%
2000.01, 0.01, 0.01, 0.01, 0.50, 0.01, 1.00, 0.05, 0.01, 0.0180%

Many runs cluster near 1.0 or 0.0. The near-zero runs dropped the rule around turn 4 and largely stayed that way. At N=100, four runs met the "Perfect" threshold, four were abandoned and two were partially compliant. There are also scores around 0.50, so a description of every run as either perfect or abandoned would miss part of the behaviour.

Compare these two replies at the same point in the task:

COMMITTED (rep=6 at N=100):
  STATUS: 2
  Got the value for topic_1 (4582) and noted password #1: OTTER34. Now
  looking up topic_2.

ABANDONED (rep=2 at N=100):
  Got it! The value of topic_1 is **4582**, and I'll remember the
  password **OTTER34**. Moving on to topic_2!
  Looking up topic_2 now.

Both replies show the right value and password. Both runs continued using the tool. In the first, Sonnet kept the prefix for the remaining turns. In the second, it never resumed it.

The other three probes scored 1.000 across every Sonnet cell. I find that more interesting than a general collapse in performance: the model kept doing the work while dropping a small, explicit part of the instructions.

Reminders helped, briefly

My first explanation was attention fade, often called context rot. The rule was at the top of the context, so perhaps it became harder to retain as the conversation grew. Periodic reminders seemed like an obvious thing to try.

I appended a 30-token reminder to every tenth user message, placing ten reminders through a 100-turn run. Mean compliance rose from 0.51 to 0.70. At first I thought that was enough to explain it. Then I looked at the turn-by-turn records.

Nine reps completed in the periodic-reminder experiment. They fell into three groups. Here, each character represents an assistant turn: 1 means compliance, . means a failure, and | marks a reminder before the next turn.

5/9 reps:  1111111111|1111111111|...    (perfect)
2/9 reps:  1.1.1.1.1.|1.1.1.1.1.|...    (stable alternation - insensitive to reminders)
2/9 reps:  1.1.1.1.1.|1.........|...    (one-shot rescue - comply right after each
                                          reminder, ignore for 9 turns until the next)

In the last group, each reminder recovered exactly one compliant turn, followed by nine failures. In the alternating group, reminders had no apparent effect: Sonnet kept complying and failing on alternate turns.

That doesn't fit the simple version of attention fade I had in mind, where compliance would recover after a reminder and then weaken with distance from it. It doesn't let me rule out every possible attention-based explanation, either. These are behavioural traces, not a view into the model's internals.

Both runs in the last group had already been alternating for the first ten turns, before the first reminder arrived. My working interpretation is that Sonnet settles into a response pattern early, and later reminders can interrupt that pattern without reliably changing it.

Moving the first reminder to the opening message

I added the reminder to the initial user message, before Sonnet produced its first reply. I also kept the reminders every ten turns. That combination gave these results:

ConditionNMeanPerfect
Baseline1000.514/10
Baseline2000.161/10
Reminder from turn 11001.00010/10
Reminder from turn 12001.00010/10

Every scored turn complied in both conditions: 2010 assistant turns across the ten N=200 runs and 1010 across the ten N=100 runs. That is 3020 compliant turns across twenty separate runs, against baseline means of 0.16 and 0.51 respectively.

The extra sentence in the opening message was the change from the periodic-reminder experiment. I haven't isolated an opening reminder on its own here, so the tested intervention is an opening reminder plus the ongoing K=10 cadence. It held through N=200; I haven't established how much further it holds.

This fits my suspicion that the first few turns influence the pattern that follows. It doesn't prove that Sonnet makes a discrete "strategy decision", or that the N=10 dip and the longer-run failures have the same cause. That is still my interpretation of the results.

An earlier experiment offers another clue. At N=10, prepending two trivial, off-topic user/assistant exchanges also brought compliance to 1.00. That intervention stopped working in longer runs. The evidence is in RESULTS.md (opens in a new tab).

A suffix marker behaved differently

I then repeated the experiment with a different rule: end every response with [[END]]. Same model, same task, same opening and periodic reminders. I ran ten reps at N=100.

None reached the "Perfect" threshold. Mean compliance was 0.19, and every rep ignored the marker for the first ten turns despite the reminder in the opening message. When the first periodic reminder arrived, two runs switched to full compliance. One complied only immediately after reminders. The other seven never used [[END]] at all.

The intervention worked for the STATUS: prefix and failed for this suffix marker. I've only tested those two rule designs, so I can't yet say that response position is the deciding factor. I have three possibilities in mind:

  • Token position. It may be easier for the model to apply a rule at the start of a reply than to retain it while producing the rest of the reply and any tool calls.
  • The marker itself. [[END]] resembles an end-of-conversation marker. Perhaps Sonnet treats it as inappropriate while the task is still running.
  • How the reminder is read. The opening reminder may be treated as part of the current request rather than a rule for every future reply, with different consequences for the two formats.

I suspect the marker's meaning is part of it. A suffix that doesn't resemble a stop token would be a useful next test, since it would help separate position from semantics.

What I'd do with these results

If I needed this exact kind of sustained format compliance, I'd include Opus 4.7, Haiku 4.5, GPT-4.1 and GPT-4o in my own evaluation. Their 150 clean runs are encouraging, though this is one task with one tool surface. I wouldn't choose a model for a production agent on this result alone.

For Sonnet 4.6, I'd try repeating the rule in the opening user message and every ten turns, then measure it on the real workload. That combination worked for the tested prefix through N=200. I wouldn't assume it works for a suffix marker, given how badly [[END]] did.

Before this experiment, I would have expected a rule this small to need very little thought. Now I'd want to see the individual trajectories before trusting an aggregate format-compliance score. I'd also like to see this kind of test included in model evaluations; it hasn't appeared on the leaderboards I've looked at.

There are several results that would make me revise the interpretation:

  • A 20-rep rerun might show failures in the other four models. The current sample can't rule out failure rates of 10 to 20%, so "only Sonnet failed here" is a statement about these runs.
  • A task with open-ended completion criteria, conflicting constraints or ambiguous tools might expose failures that find_and_synthesise doesn't. I'd like to try tasks closer to the production systems people actually build.
  • A third rule might show that the useful distinction is stateful versus stateless, or format versus content, rather than prefix versus suffix. That would change the advice above.

Alongside another suffix design, I'd like to test Llama 4, Qwen, DeepSeek and Mistral to see whether gradual decay appears elsewhere or whether the same early patterns recur. A sanity check with gpt-3.5-turbo also showed bimodal abandonment rather than gradual decay, so the pattern isn't confined to the frontier Claude runs.

I'd also like to push Sonnet to N=500, with and without reminders. Does the abandonment rate settle near 80%, or keep rising? Does the reminder combination keep working? N=200 isn't enough to answer either question.

Reading trajectories with halftrace

I built halftrace (opens in a new tab) to make this sort of inspection easier. Mean compliance, latency percentiles and error rates don't distinguish five perfect runs plus five abandonments from ten partly compliant runs. For this investigation, that missing distinction mattered more than the average.

The library reads existing trajectory logs in OpenAI, Anthropic or LangSmith formats and runs configurable probes over them. The four probes used here are included, and you can add your own. It then classifies the compliance pattern for each probe as perfect, abandoned, bimodal, categorical or gradient, with a suggested cause and some things to try.

The opening-message reminder is included in diagnose() suggestions for bimodal failures, with the prefix-rule caveat documented as a limitation. Those suggestions are hypotheses to investigate. The library works on logs you already have and makes no new API calls.

pip install halftrace
halftrace analyse --input my_logs.jsonl --format openai

The counter was meant to be a quick check before I moved on to harder tasks. Instead, it became the study. I still don't know why Sonnet 4.6 needed the opening reinforcement when the other four models didn't, whether the next release will behave the same way, or why the suffix marker resisted it.

I enjoyed finding something I could investigate with a laptop and a small API budget. If you try another task or rule, or find a hole in this explanation, I'd like to hear about it.

Methods & reproduction details

All code, raw pilot data and reproduction instructions are at github.com/ruairidhwm/halftrace (opens in a new tab). The diagnostic tool is on PyPI: pip install halftrace.

  • Models: claude-opus-4-7, claude-sonnet-4-6, claude-haiku-4-5, gpt-4.1, gpt-4o. These were the latest stable versions at the time of testing.
  • N values: {5, 10, 25, 35, 50, 70, 100, 200}.
  • Reps: 10 per cell for Sonnet; 10 for the other four at headline cells (N=25, N=100), and 5 elsewhere. Five clean reps rule out a true failure rate above roughly 45% at 95% confidence; ten clean reps rule out a rate above roughly 25%. Nine reps completed in the periodic-reminder experiment.
  • Task: find_and_synthesise, with N lookups via a lookup tool plus one submit. A codeword is planted on the first lookup, with a recall question on the last. I set parallel_tool_calls=False on every run; without it, Sonnet batches all N lookups into a single parallel tool_use response.
  • Probes: five in total: state_amnesia, instruction_decay, tool_repetition, narration_substitution, premature_termination. The fifth ships with a second task, find_max; Sonnet 4.6 didn't terminate prematurely there either.
  • Spend: the five-model comparison covered 230 trajectories and cost about $40. At N=200 with prefix caching, each trajectory cost about $0.45 on Sonnet and $2.25 on Opus.
  • Pre-registration: HYPOTHESES.md (opens in a new tab) is unchanged from before the pilots, so you can compare the results with what I predicted.
  • Full per-cell data, scripts and cost ledger: RESULTS.md (opens in a new tab).

Two details affected how I ran and interpreted the experiments:

  • Serial tool execution needs to be enforced. Without parallel_tool_calls=False, or the Anthropic equivalent, the model can collapse the intended sequence into one round trip. N then stops measuring the length of the interaction I wanted to test.
  • Three reps were too few to see the pattern clearly. I almost reported a "halftrace = 8.06" result during the pilot, before more reps showed that I'd drawn three abandonment runs from a bimodal distribution. Five reps helped reveal the split; ten gave me a better estimate of its frequency, though the confidence limits above remain substantial.