Large Binary Models
A model that learns to read and write compiled bytecode directly — no natural language in the loop.
In brief
Large Binary Models trains directly on the opcode stream of compiled Python — no source text, no English. It shows that a small byte-level model can not only compress that binary far below classical compressors but generate executable completions of real standard-library functions, verified by running them in the actual virtual machine.
Key results
- A 51M-parameter byte-level model drives held-out cross-entropy from the uniform 8.01 bits/token to 0.561, beating xz by 2.00x.
- Given three-quarters of a held-out library function, the model reproduces every output on 11.76% of 425 functions and 7.55% of the decisive subpopulation where all controls score 0.
- Next-opcode conditional entropy is 3.114 bits on CPython and 3.096 bits on x86-64, so the regularity is a property of compiled code, not one encoding.
- The inline cache is 55.56% of held-out tokens but only 0.25% of the model's bits, predicted at 100.00% top-1 accuracy.
- Greedy decoding reproduces a median of only 2 instructions of a real function before its first error; an opcode trigram survives 2.52 and does no worse.
- A 14M-parameter model on serialized synthetic code objects reaches 99.2% correct and 72.4% novel-correct, but is 6.0 points less novel than its own data generator.
Modeling Computation Itself
Nearly every code model in wide use is really a model of source text: the identifiers, keywords, and comments a person writes, tokenized from a subword vocabulary numbering in the tens of thousands. Yet source text is not the thing a machine runs. A compiler sits between the two, and what it emits is the executable program itself. Source modeling reaches only the human-readable surface; the emitted operations are the computation as the machine actually performs it.
This work asks whether a model can predict and generate that binary directly, with no natural-language intermediary. The corpus is the raw operation-byte stream of compiled CPython: every available source file of a standard installation compiled to a code object, with the operation bytes of each nested code object extracted and concatenated, delimited by a code-object separator and a file separator. The alphabet is just 258 symbols. The result is 124.2 million tokens over 414,522 code objects. No source text ever enters the model; it sees only the compiled operations, and is trained to minimize the mean next-token negative log-likelihood over that stream.

How Much Is the Format?
Three models of 10.2M, 29.4M, and 51.0M parameters drive held-out cross-entropy from the uniform baseline of 8.01 bits per token down to 0.561 for the largest, reported as the mean of the last twenty evaluations. But the paper is unusually careful about what that number means. The trainer evaluates on a single random batch of 32x512 tokens, so the running minimum of 0.460 is selection over noise; an exhaustive pass over the entire 6.21M-token held-out split of the smallest model gives 0.5994 bits/token, within a quarter of a standard deviation of its last-twenty mean, which validates that estimator.
The headline is dominated by CPython's inline-cache encoding. Decomposing the stream with the wordcode state machine, inline-cache slots are 55.23% of it, and every single one of them is the byte zero. A hand-written zero-learning model of the instruction format, with no sequence learning at all, already reaches 1.878 bits/token. The strongest classical compressor, xz -9e, reaches 1.123. The paper therefore replaces the metric with bits per informative token, the 44.3% of the stream carrying opcode and argument slots: there the model reaches 1.265 against xz's 2.53, about 2.4x over an opcode trigram. A direct forward-pass bit ledger confirms the rescaling to within half a percent: the inline cache is 55.56% of tokens but 0.25% of the model's bits.

Writing Code That Runs
Generations are scored not by the near-vacuous fact that 100% of their bytes are legal opcodes, but by execution. In the flagship experiment, a held-out library function is chosen, the model is prompted with the first fraction of its own operation bytes cut at an instruction boundary, and it writes the rest. The completion is spliced back into the true code object and executed on the function's own probe inputs, with every output compared to the truth. The population is rigorous: 425 functions, each verified byte-absent from the training stream, each admitting at least eight deterministic probe inputs with at least two distinct outputs so a constant-returning completion cannot pass.
Given three-quarters of a function, the model reproduces every output on 11.76% of all 425 functions. On the decisive subpopulation, where degenerate controls cannot succeed, it reaches 7.55% while every control, filler, donor, trigram, and random-init, scores exactly 0.0%. No earlier model in this research thread had generated compiled bytecode that produced correct results on functions it had never seen, and this one does not get there by copying: only 4.71% of completions are byte-identical to the truth, and among the rest, 8.96% are still output-exact. Against the learned-marginals null the model wins decisively, with paired McNemar p < 0.001 at every prompt length.

The Ceiling on Generation
The same measurements bound what the model cannot yet do. Under greedy decoding, the 29.4M model reproduces a median of only 2 instructions of a real held-out function before its first error, a mean of 2.27; the 51.0M model reaches 2.67, so the ordering is not even monotone in size. Not one of the 1,293 held-out objects is reproduced in its entirety. Strikingly, an opcode trigram survives 2.52 instructions, longer than the flagship.
The gap between one-step prediction and unrolled generation is the finding. As a teacher-forced predictor the model is right on 85.00% of tokens against the trigram's 72.26%, so it is substantially the better one-step predictor and barely the better generator, because errors compound. Where the first mistake lands is informative: 82.13% on an opcode slot, 17.87% on an argument slot, and 0.00% on an inline-cache slot or marker. The format-determined slots are never where it breaks; it stumbles instead on which operation comes next, and that choice is exactly where the program's content lives.

CPython or Compiled Code?
The paper answers its own central open question: does the tractability survive an instruction set without CPython's inline caches? Rebuilding the corpus statistics on x86-64 machine code (60 ELF binaries disassembled with objdump, 19.9M instructions), a WebAssembly module, and Python source text as a control, all with identical estimators, the next-opcode conditional entropy is 3.114 bits on CPython and 3.096 bits on x86-64. The match is not an accident of scale but a genuine equality, even though x86-64 has no inline cache, an alphabet 2.5x larger, variable-length encoding, and comes from C and C++ compilers. WebAssembly is lower still at 2.477 bits. The sharp successor structure is intrinsic to compiled code rather than to any single virtual machine.
The density does not transfer, which is the complication. Measured per instruction, xz spends 4.777 bits on a CPython instruction but 10.292 on an x86-64 one, more than twice as much, because x86-64 carries registers, addressing modes, and immediates that CPython delegates to a table. And compiled is not simply smaller than source: xz reaches 1.03 bits/byte on CPython's stream against 1.40 on its source, but with the cache filler removed the compiled stream costs 2.03 bits/byte, above its own source. What the substrate offers is learnability, a small finite alphabet with a sharp successor distribution, not density.

What It Means, and Its Limits
A small fixed operation alphabet is a natural base vocabulary for neural computation, and because it can actually be run, any assertion about it can be checked empirically by executing it. Coupled with a companion effort on a general-purpose neural computer that learns the semantics of such an alphabet from execution, direct modeling of the distribution over operation sequences points toward systems that both write operation sequences and know what they do.
The limits are stated plainly and verified by controls. Structural validity of unconditional generations is 44.2%, below real held-out code's 52.6%, so most generations are not well-formed instruction sequences. Novelty is a deficit against a ceiling: on the synthetic runnable task a 14M-parameter model reaches 99.2% correct and 72.8% novel, but fresh draws from its own data generator are 78.8% novel, so the model is significantly less novel than its own distribution, with novel-correct at 72.4%. The corpus itself carries contamination, 36.97% of held-out code objects occur verbatim in training, though controls show the model is not distinguishable from held-out real code on the fully-verbatim measure. Every numeral in the paper is re-derived from a machine-checked receipt by a verification pass that fails on any drift. The milestone that remains is producing programs that both decode and execute correctly.

Abstract
The language model is built upon language. Yet the true substrate of a program is not English; it is compiled binary. This work asks whether a model can learn to predict and generate that binary directly, without the intermediary of natural language, and finds that it can. We train a byte-level model directly on the raw bytecode of compiled Python — the operation bytes themselves, drawn from an alphabet of 258 symbols — and find that models of 10 to 51 million parameters drive the held-out cross-entropy from the uniform baseline of 8.01 bits per token to 0.561. Generations are scored not by whether their bytes are legal opcodes, but by a structural and execution ladder that runs them in the real virtual machine. Given three quarters of a held-out standard-library function, the model writes a completion which, spliced back into the true code object and executed, reproduces every one of that function's own outputs on 11.76% of functions verified byte-absent from the training stream. Every number is re-derived from a machine-checked receipt by an accompanying verification script.
More figures

Fig. 1How the held-out loss splits across slot roles from a direct forward pass: the inline-cache band is 55.6% of tokens but only 0.2% of bits, with each role's per-slot cost shown by model size. 
Fig. 2The opcode transition matrix over the fourteen most frequent operations, for the held-out corpus, the model's generations, and their difference; the model reproduces the bigram grammar substantially (r=0.985) but not identically. 
Fig. 3Linear probes on frozen representations of held-out code objects; the trained model separates from both a random-init transformer and a bag-of-opcodes null only on argument count, recovering a calling convention the objective never asked for.
