Model settings, prompt testing and side-by-side compare.
Prompts, model and sampling sliders, streamed output with latency, cost and run history.
Send one prompt to 2–3 models, watch staggered streams, vote best and compare stats.
Template with detected {{variables}}, live preview and a test-case table with pass/fail.
Score trend across runs, KPIs vs baseline and a per-example table with word diffs.