Prompt/Model Benchmarking Tool
notebook-ta ships a local GUI tool that helps instructors iterate on system prompts and compare
LLM/model combinations before rolling them out to students. It is documented in detail in
the functional spec and
the architecture spec
(architecture).
Launching
The benchmarking application is an optional instructor tool. Install its extra before launching
it; the base notebook-ta install used by students does not include the NiceGUI stack.
pip install "notebook-ta[bench]"
notebook-ta bench # show the project welcome screen
notebook-ta bench my_project.json # offer this project on the welcome screen
This opens the benchmarking UI in your default browser. The welcome dialog lets you reopen the
most recent project, browse for another project, or create a new one. New projects require a name
and an exercises TOML file; the name becomes the suggested JSON filename on first save. Their tag
list starts with correct, wrong complexity, logic flow, and missing edge-case.
The server listens on 127.0.0.1 and has no authentication. It is intended for one trusted local
operator and must not be exposed through a proxy, tunnel, shared host, or non-loopback bind.
Editable Python runs in timeout-bounded workers with the host user’s permissions; these workers
are not security sandboxes. Read the trust and security model before using
third-party projects or student submissions.
Workflow
Settings — configure the internal model (used only to help draft example student solutions—never to score benchmark output), global setup code, Python paths, tags, and autosave. Global setup code runs for every Python exercise before that exercise’s own setup code; both blocks share the solution namespace. Every tag has an editable color which is used for its badges throughout the app. Save As opens a native file picker. Close project returns to the welcome dialog; unsaved changes require explicit confirmation before they are discarded.
Exercises — exercises are expanded by default and their solution cards are arranged side by side with horizontal scrolling. Edit exercise and solution display names inline, append new exercises to a local TOML catalog, add solutions manually or with the internal model, tag them (e.g.
correct,wrong complexity), and run their unit tests. For each exercise, benchmark-only setup code can define helper variables or functions before unit tests run; this setup is saved in the benchmark project JSON file, not inexercises.toml. Exercise edits preserve the catalog’s comments and formatting; remote TOML catalogs are read-only. Free-text exercises use a prose editor, hide setup/test controls, and can generate tagged draft prose with the internal model. Their answers are never executed.Runner — write the
on_success,on_failure, and free-text evaluation prompts to test, select one or more models, and click Run Benchmark. Prompts are frozen into a versioned snapshot (V1,V2, …) the moment you click Run, so past results always remain reproducible even if you keep editing the prompt afterward.Compare — review results in a matrix whose rows are exercise/student-solution pairs and whose columns are historical model + prompt-version combinations. The latest run is selected by default; use the shared multi-select to compare other combinations and the tag filter to narrow the solution rows. Column headers show average TTFT, total generation time, and throughput for the visible results. The left column shows each Python solution or free-text answer and opens the full exercise statement. Click any result cell to inspect its exact prompt, unit tests, metrics, and errors. Runs can be permanently deleted from this tab after acknowledging that deletion cannot be undone; all results produced by that run are deleted with it. If the exercise or solution changed after generation, the cell is flagged ⚠️ Stale (Inputs Modified) with a one-click Re-run.
API credentials
The benchmark UI never accepts or stores an API key value. For the internal model and each model under test, enter the name of an environment variable in the API key environment variable field. Set that variable before launching the benchmark application. For example:
$env:NOTEBOOK_TA_OPENAI_KEY = "your-api-key"
notebook-ta bench
export NOTEBOOK_TA_OPENAI_KEY="your-api-key"
notebook-ta bench
Enter NOTEBOOK_TA_OPENAI_KEY—not the key value itself—in the UI. The application resolves the
variable only when it creates the LLM provider. Local providers such as Ollama normally leave this
field empty.
Only the environment-variable name is saved in the benchmark project. The resolved secret is kept in memory and is never written to project JSON. Ensure the variable is available in the environment whenever you reopen the project and run that model.
Project files
Everything (settings, global and per-exercise setup code, student solutions, prompt version history, and the full execution history with metrics) is saved to a single JSON project file via the Save button or autosave. API key values are excluded; project files contain only their environment-variable references. The autosave interval must be a positive number of seconds.
Existing schema-v2 project files remain compatible; projects without global setup code load with an
empty default. The persisted code and student_code field names are retained for both answer
types; snapshots additionally record the answer type, evaluation criteria, and both setup blocks.