Where answers come from

Some pages can ask a language model questions. Set this up once; your choice is remembered in this browser for the whole site.

Model source
Use another source instead

Smallest-to-largest weight downloads. Qwen3.5 0.8B is our starting choice for page Q&A. GPU estimates are not total RAM: leave room for the browser, context and temporary buffers. Hardware changes speed and capacity; answers still need checking.

Session-only loading avoids saving model weights, uses extra RAM and downloads again after unloading. Turn it off to save downloads. Delete-data-on-exit or private browsing may impose a much smaller storage limit than the displayed estimate.

You sign in on OpenRouter; it gives this page a key for your account.

A free account does not make every model free. Choose a model with a listed :free variant; merely adding that suffix does not create one. Check free-model availability and pricing.

Chat

Inspect the page excerpt

Up to 4,800 characters, not the whole site. Recent conversation turns are retained within the model’s context limit.

Ask about this page, or just say hello.

Ready when you are.

Not saved or uploaded automatically. Closing keeps this conversation in this tab; changing pages starts a new one. Model replies can be wrong.

Labs / model comparison

One opening.
Three different continuations.

The same three words, different learned patterns. Compare what these models write—and the alternatives they could have chosen.

The recorded Chrome trial reproduces three Session 3 observations from October 8, 2026, with corrected odds. These are generated completions, not claims about the soul or evidence that a model holds beliefs.

Back to the course experiment →

Try the same opening

No model downloads until you press Compare. Runs load one model at a time—about 1.1 GB total weights for this trio, plus runtime files. Your saved site model choice is unchanged.

Downloads, memory and privacy

Session-only mode uses extra RAM and downloads again on each run. Saved downloads can fail despite an ample reported quota. Other selections change download and memory needs; estimates appear in each picker. Models unload between trials and after completion or cancellation.

Prompts stay in this tab; downloads contact the model hosts. Results last only for this visit; export before leaving. Exports include your opening words and generated text.

Earlier observations below. Choose your models and run a fresh comparison whenever you’re ready.

Run a comparison to inspect the odds.

At token 1, all models see exactly the same opening. Later, each follows its own continuation. Tokens aren’t always whole words. ↵ marks one or more newlines; exact whitespace is preserved in the export.

SmolLM2 360M

The soul is the seat of the divine spark within us

Earlier observation · October 8, 2026 · transcribed from a Session 3 screenshot

Completion text only. No usable odds were preserved from this earlier trial.

Qwen3 0.6B

The soul is a concept that is both present and absent

Earlier observation · October 8, 2026 · transcribed from a Session 3 screenshot

Completion text only. No usable odds were preserved from this earlier trial.

Gemma 3 1B

The soul is is soul ↵ soul ↵ soul ↵ soul

Earlier observation · October 8, 2026 · transcribed from a Session 3 screenshot

Completion text only. No usable odds were preserved from this earlier trial.

What are we measuring—and what changed?

Fresh runs use raw text completion, not a chat template. They measure the five strongest next-token probabilities at temperature 1, without repetition or frequency penalties, then continue with the strongest candidate for up to eight tokens. A recognized end token stops the continuation. Different tokenizers and quantized builds mean this is an experiment, not a model quality ranking.

Gemma’s repetition is a known issue with this build/runtime combination, not a judgment about all Gemma models. All six builds remain available in each picker. Your hardware, browser and network affect loading and speed.

The earlier screenshots requested odds at temperature 0. In this WebLLM runtime, that collapses the distribution before log probabilities are returned, producing misleading 100% bars and negligible alternatives. We preserved the completion text but deliberately did not copy those bars. Runtime probability calculation.

Fresh bars show the reported probability, not a percentage renormalized among only five candidates. The remaining probability is shown as “Other tokens.” Seed 42, pinned model revisions and runtime details are recorded in the export. Timings depend on your hardware, network and browser cache.