Skip to content
SLM-125M

A 125M legal & financial language model

Trained from scratch on US court opinions, SEC filings and educational web text. Write the start of a sentence and watch it continue.

Start a sentence

Connecting…

… tokens of 1024 context · 33/4000 characters

Legal

SEC filings

General— Its weak spot, on purpose: watch it stay fluent while getting facts wrong.

Completion

Pick a sample or write your own opening, then press Complete. The model continues your text; it doesn't answer questions.

The model, in numbers

Parameters
125,848,320
all trainable
Held-out loss
2.127
nats per token
Perplexity
8.39
on held-out text
Bits per byte
0.657
tokenizer-independent
Tokens
8.16B seen
2.04B unique × 4 epochs
Context · vocab
1,024 · 16,384
tokens · BPE vocabulary

Pretrained from scratch on US case law, SEC filings and educational web text. See exactly how it was built →

How surprising is your sentence?

Perplexity, made personal: the model scores every token of your text by how much it expected it. Familiar legal and financial phrasing scores low; nonsense scores high.

Scores appear here. The first request after idle takes ~12 s while the model wakes.