Marc's Blog

About Me

My name is Marc Brooker. I like to build things that work, and do cool stuff. I like building big things. I also dabble in machining, welding, cooking, and skiing.

I am an engineer at Amazon Web Services (AWS) in Seattle, where I work on agentic AI, especially safety and policy for agentic AI. Before that, I worked on EC2, EBS, databases, serverless, and serverless databases.
All opinions are my own.

Links

My Publications and Videos
@marcbrooker on Mastodon @MarcJBrooker on Twitter


Is this blog written by AI?

Strands Decider: Why Not an Encoder?

Expertise level: exhausted.

As I admitted in my last post, I am very much not an expert model developer. What follows is likely to be, at least partially, inaccurate. I have tried to be careful and quantitative, to partially balance a lack of deep expertise.

Since we launched strands-decider-2B last week, a couple of people have asked me why it’s a modified LLM and not an encoder (like BERT). Their instinct seems to be that bidirectional (i.e. each token’s representation is built from tokens before and after it) is likely to out-perform causal (i.e. only before tokens are used). This is the same reason encoders are widely used as classifiers.

This is also a fair question, because if you squint at it right, we’re abusing a decoder as an encoder.

To test this hypothesis, I compared five designs:

I also measured frozen ModernBERT and ettin-encoder-1b, but found that they under-performed T5Gemma parameter-for-parameter, and so didn’t invest time in training them.

First, let’s look at accuracy and calibration.

b2 does very well here, beating v19 on JevBench accuracy and Brier, and on JF100 Brier at the same accuracy. This is good, but recall that b2 has ~8B total parameters compared to v19’s 2B, and so is not an apples-to-apples comparison.

The really interesting comparison is v19 and e1b. e1b slightly beats v19 on JevBench accuracy (but within the run-to-run variance on v19), with slightly worse calibration, and ties on calibration on JF100 with slightly weaker performance.

So far, not a slam-dunk for the encoders.

Next, let’s turn to latency.

The encoder-only models win handily on latency, with a median of around 35ms compared to v19’s 65ms. The encoder-decoder models are slowest. With half the latency and similar performance, this is looking like a big win for e1b. But thing get a little more complicated when we dive into the data.

Here, we see that e1b wins handily at lower prompt lengths, but scales worse as the prompt grows. This is consistent with the great median performance. Whether this matters depends on the workload mix: both JF100 and JevBench trend small, to e1b’s benefit. The scaling difference here comes down to the compute costs of the two models: e1b has 26 full attention layers with the associated quadratic compute cost, while 75% of v19’s layers are Gated DeltaNet with linear compute cost. These layers have a higher fixed cost, but better asymptotic cost, leading to the higher latency floor2.

Finally, let’s look at one interesting sub-plot, going back to e1a (our unmasked 2.6B encoder-only candidate). On the eval sets, it beats v19 on seven of nine tasks, and is very slightly behind on two. But it does significantly worse at JevBench (153 vs 168), and a little worse at JF100 (158 vs 162). This suggests that the encoder generalizes worse, at least without the masked approach e1b takes. Is being able to attend to the question when reading the state actually a weakness?

Evaluation set v19 e1a Δ
Held-out short tasks 0.647 0.695 +0.048
MuSiQue 0.882 0.948 +0.066
ContractNLI 0.873 0.864 −0.009
BoardgameQA 0.820 0.884 +0.064
HotpotQA (held out) 0.719 0.748 +0.029
Generated, v16’s 0.857 0.851 −0.006
Generated, v18’s 0.777 0.786 +0.009
Adequacy, HelpSteer2 0.722 0.761 +0.039
Adequacy, generated, balanced 0.788 0.821 +0.033

Where does that leave us on the encoders vs decoders question? Not very much closer to the truth, I fear. This experiment is hopelessly confounded, changing architectures, base models, instruction tuning vs base, and other variables. But we have learned some interesting things. Neither seems like a slam-dunk winner for this kind of work at this size.

The code for this experiment is in a branch for now. It doesn’t seem worth pulling into mainline strands-decider yet, but may with some development.

Footnotes

  1. Some decoder ideas sneaking back in. Decoders do the same trick (KV cache and prefix cache) for the same reasons. Certain not for the first time memoizing incremental results has helped performance.
  2. Certainly not the first time an $O(N^2)$ algorithm has beaten an $O(N)$ one at small data sizes.