Where answers come from

Some pages can ask a language model questions. Set this up once; your choice is remembered in this browser for the whole site.

Model source
Use another source instead

Smallest-to-largest weight downloads. Qwen3.5 0.8B is our starting choice for page Q&A. GPU estimates are not total RAM: leave room for the browser, context and temporary buffers. Hardware changes speed and capacity; answers still need checking.

Session-only loading avoids saving model weights, uses extra RAM and downloads again after unloading. Turn it off to save downloads. Delete-data-on-exit or private browsing may impose a much smaller storage limit than the displayed estimate.

You sign in on OpenRouter; it gives this page a key for your account.

A free account does not make every model free. Choose a model with a listed :free variant; merely adding that suffix does not create one. Check free-model availability and pricing.

Lab / browser inference

Small models. Real questions.

Say hello, then ask a model to read a page with you. Six experimental builds, one shared test. Your hardware is part of the experiment.

Qwen3.5 0.8B is our starting choice for page Q&A—not a guarantee of accuracy. Models are listed smallest-to-largest by weight download; estimated GPU memory is shown separately. RAM, GPU, browser settings and context length affect whether a model loads and how it performs.

Open the chat pane ↓

Meet the models

Observed on 2026-10-08: Chrome 155, Apple / Metal 3, 16 GiB memory. Session-only loading, WebLLM 0.2.85, temperature 0, seed 42, Qwen thinking off. Reply timings are medians across four questions about one article. They include incorrect answers: speed is not an accuracy score.

How to interpret these numbers

This was a smoke test, not a benchmark. We checked a greeting, name recall, three page questions and one question with no answer in the page. We then challenged the two Qwen finalists with matched follow-ups. Neither was consistently reliable. No overall score is claimed: simple keyword checks incorrectly counted some wrong answers as correct.

Download sizes count quantized weights, excluding runtime and tokenizer. GPU memory estimates come from the WebLLM catalog at a 4,096-token context, not measured total RAM. The browser, temporary buffers and session-only weight storage need additional memory. Gemma’s estimate also predates our window override. Leave headroom; a listed size is not a promise that it will fit.

Load times are observed download/reload times, not controlled cold-start benchmarks. Use this pane to measure your own machine and inspect the actual answers before trusting a result.

Try a conversation

Type a message and press Send. If needed, the model downloads first and your message sends automatically when it is ready. You can also preload with Load model or run the smoke test. Session-only mode uses extra RAM and downloads again after unloading; saved downloads may hit browser storage limits even when the displayed quota looks ample.

Preparing the lab…

Sample: A by-line for every model. Changing context starts a new conversation. Prompts stay in this tab; downloads contact the model hosts. Nothing is sent to an inference service by this Lab.

Inspect exactly what the model receives
Loading the public article…

Conversations and results are held only for this visit, not saved automatically. Leaving the Lab unloads its model and clears them; export first if you want to keep them. Exports contain your prompts and generated replies. Never treat a model reply as verified source material.

Smoke-test replies and review guide
No tests run this visit.

Loading your model

Your message is waiting. It will send automatically when the model is ready. The first download may take a little while.

Preparing the model…