Skip to content
SLM-125M

How it was built

From public data to a working legal language model

Every step, what it does, and why, with the real numbers from the run. Nothing here is a mock-up: the charts are built from the training logs and pipeline reports.

  1. 13.6B

    clean tokens available

    measured across 3 public datasets

  2. 2.04B

    unique tokens kept

    after cleaning, dedup, decontamination

  3. 8.16B

    tokens trained on

    4 epochs

  4. 125.8M

    parameters

    Llama-style, from scratch

  5. ≤ $25.32

    of compute

    44.9 min on 8 × H100

Chapter 1

The data budget problem

The obvious recipe for a legal model is mostly legal text: say 70% court opinions, 20% SEC filings, 10% web, about 10 billion tokens in all. Before downloading anything we measured how much clean text each source actually holds, by cleaning 2,000 sample documents per source and projecting to the full dataset.

The case-law dataset has only 282,390 documents. After cleaning that is about 0.81B tokens, while 70% of 10B would need 7.00B: a 8.6× shortfall. The recipe was impossible.

Demand vs supply, in tokens

Outline: what a 70/20/10 mix of 10B tokens would need. Solid: clean tokens the source actually contains (measured).

  • Available (measured)
  • Needed for 70/20/10

US case law

Needed
7.00B
Available
0.81B

SEC filings

Needed
2.00B
Available
1.16B

Educational web

Needed
1.00B
Available
11.67B

The mix we actually trained on

Training tokens after cleaning, dedup, decontamination and tokenization: 2.04B in total.

  • US case law35.1%715.3M
  • SEC filings42.2%860.6M
  • Educational web22.8%464.2M

Budgets are counted in proxy tokens (characters ÷ 4) during cleaning. Case law and web text hit their caps (1.0B and 0.5B); SEC ran out of documents before its 1.3B. Dedup, decontamination and the real tokenizer then shrank every source to the training totals above.

Takeaway — Measure supply before you plan the mix. The two legal sources hold only ~2B clean tokens, so we use all of them and cap the web: 77% legal, not 70/20/10.

Chapter 2

Streaming, not downloading

The three datasets live on Hugging Face as parquet shards. We never download them. Each shard gets its own container that streams rows, cleans each document on the fly, writes the survivors, and stops at its token cap. In total, 718,780 documents were streamed out of more than 10.0M available.

Why one container per shard? Cloud machines get preempted, and a preempted job restarts from zero. With one worker per shard, a preemption costs one shard, not the whole run. That happened for real: twice during tokenization, and only the affected shard re-ran.

20 workers, one per shard

Each tile is one container: documents streamed → kept, and how full it got against its token cap (case law 100M, SEC 260M, web 100M). The web dataset has 14 shards; we only needed 5.

US case law · 10 parquet shards

  • shard-000

    28,532 → 26,786

    100.0M tok · hit cap

  • shard-001

    21,535 → 21,289

    100.0M tok · hit cap

  • shard-002

    24,278 → 23,086

    100.1M tok · hit cap

  • shard-003

    33,009 → 31,777

    100.0M tok · hit cap

  • shard-004

    23,154 → 23,136

    100.0M tok · hit cap

  • shard-005

    20,815 → 20,673

    100.0M tok · hit cap

  • shard-006

    19,970 → 19,799

    100.0M tok · hit cap

  • shard-007

    21,334 → 21,242

    100.0M tok · hit cap

  • shard-008

    22,141 → 21,138

    100.0M tok · hit cap

  • shard-009

    23,439 → 23,366

    97.7M tok · source ran out

SEC filings · 5 parquet shards

  • shard-000

    9,068 → 8,928

    195.4M tok · source ran out

  • shard-001

    9,487 → 9,372

    216.3M tok · source ran out

  • shard-002

    10,106 → 10,003

    251.7M tok · source ran out

  • shard-003

    9,719 → 9,630

    260.0M tok · hit cap

  • shard-004

    9,372 → 9,266

    260.0M tok · hit cap

Educational web · 14 parquet shards

  • shard-000

    85,747 → 82,894

    100.0M tok · hit cap

  • shard-001

    86,863 → 83,872

    100.0M tok · hit cap

  • shard-002

    85,351 → 82,548

    100.0M tok · hit cap

  • shard-003

    87,383 → 84,458

    100.0M tok · hit cap

  • shard-004

    87,477 → 84,695

    100.0M tok · hit cap

+ 9 more shards not needed

Proxy tokens = characters ÷ 4 (the real tokenizer didn't exist yet).

Takeaway — Stream, don't hoard, and fan out one worker per shard so that a failure costs a shard, not the run.

Chapter 3

Cleaning

Raw text is noisy: court opinions are often scanned paper with OCR errors, and filings are full of headers and signature blocks. Every document passes through the same six fixed rules, cheapest check first. Every rule is deterministic, so the same input always gives the same corpus, and every drop is counted by reason.

Of 718,780 documents streamed, 670,124 made it into the final corpus (93.2%).

The document funnel, from stream to corpus

Phase 1 (cleaning) then Phase 2 (dedup + decontamination). Bars share one scale, starting at zero.

Streamed
718,780
− too short after line filters
698,649
− OCR garble (case law only)
697,964
− not English
697,958
− near-duplicates (MinHash)
696,352
− exact duplicates
694,301
− benchmark contamination
670,124
The six rules, in order
#StepRuleWhat it catches
1Line filterdrop lines under 40 characters, or over 30% non-alphanumericpage numbers, stamps, OCR fragments, table rules
2Boilerplatedrop lines matching known patternsFORM 10-K headers, “Page 3 of 9”, /s/ signatures, “All rights reserved”
3Length gatedrop the document if under 600 characters surviveone-line orders, stubs
4Repetitiondrop if the 10 most common 4-word phrases cover over 50% of the textspam, tables flattened into text
5LanguageASCII ratio first; langdetect only for the 90–99% ambiguous bandnon-English text, cheaply
6OCR gatecase law only: drop if over 20% of words are not in a dictionarybadly scanned opinions

Real documents, before and after

Fetched from the public datasets and run through the actual cleaning code. Struck-through lines are removed. Names of private parties are redacted.

✓ KEPT12,724 → 10,729 characters · 70/254 lines removed

Kept. Caption debris, stamps and OCR fragments are cut line by line; the opinion survives.

Show what each rule removes
  1. LAW LIBRARYtoo short
  2. on ILICATION. i ootoo short
  3. IN THE SUPREME COURT OF THE STATE OF HAWAT'T
  4. ---000-too short
  5. [NAME], Plaintiff,too short
  6. [NAME],too short
  7. the State of Hawai'i and [NAME]:too short
  8. omgtoo short
  9. No.too short
  10. 29477too short
  11. ORIGINAL PROCEEDING stoo short
  12. DECEMBER 16, 2008too short
  13. MOON, C.J., LEVINSON, NAKAYAMA, ACOBA, AND DUFFY, JJ.mostly symbols
  14. Pex Curiam. In this original proceeding, plaintiff
  15. [NAME], an unsuccessful congressional candidate in the
  16. November 4, 2008 general election, challenged, pursuant to
  17. Hawai'i Revised Statutes (HRS) § 11-172 (1993)! and HRS § 11-mostly symbols
  18. 174.5 (Supp. 2007)%, the election results for the first
  19. HRS § 11-172 (1993) Contests for cause; generally.
  20. With respect to any election, any candidate, or qualified
  21. political party directly interested, or any thirty voters of
  22. ‘any election district, may file @ complaint in the supreme
  23. court. The complaint shall set forth any cause or causes,
  24. such as but not limited to, provable fraud, overages, or

These excerpts come from the start of each dataset, where case law is unusually noisy. Across the whole run only 685 of 238,207 case-law documents (0.3%) failed the OCR gate.

Takeaway — Cheap, deterministic rules, counted by reason. Most of the work happens inside documents, where line filters cut the noise but keep the text. Only 2.9% of documents were rejected outright by cleaning.

Chapter 4

Duplicates & MinHash

Courts reuse templates, and the same filing can appear twice. Duplicates waste training budget and invite memorization. Exact copies are easy: hash the normalised text. Near copies, the same order with a different date, need MinHash.

Each document becomes a set of overlapping 5-word shingles. Two documents are near-duplicates when their sets overlap heavily: Jaccard similarity of at least 0.8. Comparing every pair of 232,292 opinions is 27 billion comparisons. MinHash compresses each document to 32 numbers whose agreement estimates the Jaccard similarity, and LSH buckets them so only likely pairs are ever compared.

Try it: two court orders (illustration). Edit either one.

Change a word and watch the shingles, the exact Jaccard, and the MinHash estimates move.

Shingles A / B
77 / 77
74 shared
Exact Jaccard
0.925
≥ 0.8: near-duplicate
MinHash estimate
0.938 · 0.906
32 · 128 permutations
Our LSH flags it
84%
probability (3 bands × 10 rows)

How likely LSH is to flag a pair, by true similarity

Our setting (32 permutations, 3 bands × 10 rows) catches 29% of pairs at Jaccard 0.8 and 72% at 0.9. With 128 permutations the curve is steeper.

  • Ours: 32 permutations
  • 128 permutations (comparison)

Exact duplicates (same normalised text) were removed separately: 1,989 SEC filings and 62 web pages. Exact dedup runs within each shard, so an identical document in two different shards can survive; this is a known, accepted limit.

Takeaway — MinHash + LSH finds near-copies in one pass. With 32 permutations it is conservative: it removed 1,606 near-duplicate opinions, but it catches only 29% of pairs at exactly 0.8 similarity, so that count is a floor, not a full census.

Chapter 5

Decontamination

We want to test the model honestly on CaseHOLD, a legal benchmark built from court opinions. If its test passages also appear in the training text, a good score would just mean memorized, not learned.

So we collected every 13-word window of the benchmark (480,908 of them) and dropped any training document that shares even one. Thirteen words is long enough that a shared window almost never happens by chance.

Try it: does this document share a 13-word window with the benchmark?

The benchmark passage is a real CaseHOLD test item; the document is an illustration. Edit either one.

Training document, with shared 13-word windows highlighted:

In valuing the remaining parcel, the trial court relied on the owners' appraiser. As other courts have explained, fair market value is not to be determined in a rarefied realm of abstract calculation, but from the perspective of a hypothetical buyer. We therefore affirm the award of severance damages.

✕ DROPPED11 shared 13-word windows: one is enough to drop the whole document.

Documents removed as benchmark-contaminated

Share of each source's documents entering the dedup phase. Web text was not checked: CaseHOLD comes from court opinions.

US case law
24,002 · 10.3%
SEC filings
175 · 0.4%

Takeaway — Strip benchmark text out of training, or the evaluation lies. This was strict: it removed 10.3% of case law to keep the test honest.

Chapter 6

The tokenizer

Models read tokens, not characters. We trained a fresh byte-level BPE tokenizer on this corpus: it starts from the 256 possible bytes and repeatedly merges the most frequent adjacent pair, until it has 16,384 tokens. Because it was trained on legal and financial text, words like “plaintiff” and “pursuant” became single tokens.

Byte-level means there is no unknown token, ever. Any string, including OCR garbage, emoji or§, can be encoded.

Try it: the real tokenizer, running in your browser

Same vocabulary and merges as the model, verified token-for-token against the Python tokenizer. A · marks a space; byte-level BPE glues the leading space onto the word.

Loading the tokenizer…

Characters per token, by source

More characters per token = the tokenizer compresses that kind of text better. Measured on 500 documents per source.

US case law
4.13
SEC filings
4.98
Educational web
4.23

Why 16,384 tokens?

Every vocabulary entry is a row of 768 numbers. At 16K that is 12.6M parameters (10.0% of the model). GPT-2's 50,257 would need 38.6M (25.4% of the same model), budget taken from the layers that actually learn language.

Reserved special tokens

  • 0 <|bos|>
  • 1 <|eos|>
  • 2 <|pad|>
  • 3 <|unk|>

Every document ends with <|eos|>, so documents in training always start right after one. That is why prompts are prefixed with <|eos|>, never <|bos|>, which never appears in the training data at all.

Takeaway — A small, domain-trained vocabulary is efficient on this domain and cheap in parameters. 16K tokens cost 10% of the model; GPT-2's 50K vocabulary would cost 25%.

Chapter 7

Packing & the honest split

Training needs fixed-size inputs. Every document is tokenized, followed by an <|eos|>, and the whole stream is cut into 1,024-token windows, with no padding wasted. Each window is stored as raw 16-bit integers: a 16,384-token vocabulary fits in 16 bits, half the disk of 32-bit.

1% must be held out for validation, and how you choose it matters. Our first version followed the original spec: every 100th window. But documents span windows. An SEC filing averages 18.8 windows, so a validation window's other ~18 windows sat in training. After 4 epochs over those siblings, “held-out” loss would partly measure memory.

Packing 12 documents into 1,024-token windows (illustration)

Each colour band is one document; ▍ marks its <|eos|>. Hatched windows are validation.

  1. train 1
  2. train 2
  3. train 3
  4. val 4
  5. train 5
  6. train 6
  7. train 7
  8. val 8
  9. train 9
  10. train 10
  11. train 11
  12. val 12
  13. train 13

▲ 5 documents leak:part of each sits in a validation window while the rest of it (tinted red) is in training. The model trains on the rest of the very text it is tested on.

How many windows a document spans

  • US case law3.41 windows
  • SEC filings18.84 windows
  • Educational web1.09 windows

The longer the documents, the worse a window split leaks.

The final split (document-level)

Train tokens
2.04B
Validation tokens
19.6M
Train documents
663,486
Validation documents
6,638
Window
1024 × uint16

A document is validation when blake2b(text) mod 100 = 0: stable on every machine, so duplicates land on the same side too.

Takeaway — Split by document, not by window. We rebuilt the split so that 6,638 whole documents are held out and every document is on exactly one side. It costs nothing and makes the headline number honest.

Chapter 8

Architecture

A Llama-style decoder: 12 identical layers of attention and MLP, 768-wide, with 12 heads, RMSNorm, rotary positions (RoPE), SwiGLU and tied input/output embeddings. These are the same ingredients as much larger open models, so everything here scales up.

Exactly 125,848,320 parameters, all trainable. Click a block to see what it does.

  1. Decoder layer × 12 · 9.4M params each

SwiGLU MLP

768 → 3072 (gate, up) → 768 (down)

7,077,888 parameters per layer · 84,934,656 in all 12

Where most of the knowledge lives: two-thirds of all parameters. A gated feed-forward network applied to every token independently.

Where the parameters are

MLP (SwiGLU)
84.9M · 67.5%
Attention
28.3M · 22.5%
Embedding (tied)
12.6M · 10.0%
Norms
19.2K · 0.02%

Capacity explorer: resize the model

Same formula as the training config. The MLP is kept at 4× the width.

Parameters
125.8M
this model
Embedding share
10.0%
12.6M in the vocabulary table
Tokens per parameter
64.8
at our 8.16B tokens seen
Chinchilla-optimal data
2.52B
≈ 20 tokens per parameter

Takeaway — Two-thirds of a transformer's parameters are its MLPs. A small vocabulary and tied embeddings spend the budget there.

Chapter 9

Training

15,564 optimizer steps of 524,288 tokens: 8.16B tokens, four passes over the corpus, on 8 × H100 in 44.9 minutes. The learning rate follows one continuous schedule: 381 warmup steps to 6e-4, then a single cosine decay to 6e-5, never restarted, even across epochs.

Repeating data risks memorization, so the held-out curve (documents the model never sees) is the one that matters. It improved at every one of its 16 evaluations.

Loss over 8.16B tokens: training vs held-out

The first ~0.25B tokens (warmup, loss 10.4 → 3.3) are above the top of the axis. Hover or use ←/→ to read values.

  • Train (500-step mean)
  • Held-out validation

Learning rate (same x-axis): warmup, then one cosine to the floor

Was each extra epoch worth it?

Held-out bits/byte

0.6572

Gain over previous epoch

−0.0132

Perplexity

8.39

Measured at the last evaluation inside each epoch (step 15,564). Each later epoch bought about half the previous gain.

The one change that made 4 epochs affordable

Measured on 1 × H100 before the real run. Cost = projected cost of all 4 epochs at that efficiency, against the $50.00 cap.

Plain PyTorch
$51.67 ✕ over cap
torch.compile
$24.28 ✓

In a model this narrow, the many small operations around each matrix multiply (normalisation, activations, rotary positions) dominate. torch.compile fuses them. On 8 GPUs the real run then reached 3.4M tokens/s at 32.0% efficiency: near-perfect scaling.

Takeaway — Four epochs helped but with halving returns. Training loss steps down at each new pass (partial memorization), while held-out loss keeps falling smoothly. That is no overfitting, but the corpus is close to exhausted at this size.

Chapter 10

Results, limits & cost

On held-out documents the model reaches 0.657 bits per byte (loss 2.127, perplexity 8.39). Bits per byte is the honest headline: it doesn't depend on the tokenizer, so it compares across models. MMLU-style benchmarks are near-random at this size.

The split by source shows what it learned: formulaic SEC language best, case law well, general web text worst.

Held-out bits per byte, by source (lower is better)

96 held-out windows per source, each scored with its own bytes-per-token.

SEC filings
0.463
US case law
0.687
Educational web
1.017

Real completions (prompt in grey)

  • The court held that the defendant’s failure to file a statement of claim for the $5,000 bond with the plaintiff’s counsel on or before February 9, 2003, did not prevent the plaintiff from filing a lawsuit against the defendant. The court reasoned that the defendant’s failure to file a claim for the bond with

  • The Company's net revenues for the fiscal year ended January 31, 1998 were $298,384,000 as compared to net revenues of $262,906,000 for the fiscal year ended January 31, 1997 ("fiscal 1997"). The net loss for fiscal 1997 was $35,917,000 as compared to net

  • Pursuant to Section 10(b) of the Securities Exchange Act of 1934, the Registrant has duly caused this report to be signed on its behalf by the undersigned, thereunto duly authorized.

    ✓ word-perfect 10-K boilerplate

  • Photosynthesis is the process by which plants produce more oxygen, carbon dioxide and water from the sun.

    ✕ fluent but false: the model's general-knowledge limit

What it cost

Data, cleaning, dedup, tokenizer, tokens (CPU)$0.00
Calibration (1 × H100, 8 min)≤ $0.54
Pretraining (8 × H100, 45 min)≤ $24.78
Total≤ $25.32

Upper bounds from container lifetimes. Serving: Modal CPU, scale-to-zero: $0 while idle.

Limits, plainly

  • A base model: it continues text and does not follow instructions or chat.
  • Invents facts and citations with total confidence. Not legal or financial advice.
  • May reproduce text from public court opinions and SEC filings, which name real people and companies.
  • Skewed to US law and to 1990s-era 10-K filings; case law is partly OCR'd.
  • No safety, instruction or preference tuning.

Takeaway — It speaks the domain fluently and does not know facts. That is exactly what a 125M base model trained on this data should do.