Stage 2 exercise: Change one thing, then check the score¶
This exercise does one thing: have two prompts answer the same six questions, then compare their scores.
You will follow this short path:
same six questions → run the original → add three examples → run again → compare scores
Step 1: Run the no-model version first¶
Run this in the folder:
python starter.py
You will see 3/6 for the original and 6/6 after adding examples. These are fixed answers built into the program to teach the workflow; this is not a model leaderboard and does not prove that examples will improve the score every time.
Step 2: Confirm that the program calculates correctly¶
python test.py
python test_anthropic.py
Neither test needs an API key or connects to a model. Seeing 4/4 passed and 2/2 passed means you are done.
🎓 Learning mode: First run the provided
starter.py(python starter.py), then change exactly one small thing and run the existing tests again:python test.pyandpython test_anthropic.py. If a test fails, undo or fix that one change and try again. You do not need to rename the file or rewrite the whole solution. Seedocs/HOW_TO_USE.mdfor the full method.
Optional: run a real model with local Ollama (Path A)
pip install -r requirements.txt
ollama pull gemma4:e4b
ollama serve
python starter.py --live
The program calls the local model 12 times: six questions with the original prompt, then the same six with the improved prompt. API cost is $0, but it uses your computer's time and electricity. A small model's score may differ from run to run, which is exactly why fixed questions and repeated tests matter.
Optional: run a real model with Anthropic (Path B)
pip install -r requirements.txt
export ANTHROPIC_API_KEY=sk-ant-...
python starter_anthropic.py --live
On Windows PowerShell, use:
$env:ANTHROPIC_API_KEY = "sk-ant-..."
python starter_anthropic.py --live
The default model is claude-haiku-4-5. One short prompt call is estimated below $0.001; 12 calls are estimated below $0.01. Actual cost depends on token count and current official pricing. Set a $0.05 total cap for your first run; see Anthropic's official pricing.
How the program works, common snags, and further reading
| Part | Plain-language explanation |
|---|---|
CASES |
Six fixed test questions, each with a correct label |
build_prompt() |
The original and improved versions differ by three examples only |
evaluate() |
One point for each correct answer |
--live |
Replaces built-in answers with real model answers |
Common snags:
- An answer such as
billing, because...is marked wrong because the output rule requires one label only. - If Ollama cannot connect, make sure
ollama serveis still running. - For an Anthropic authentication error, check the environment variable; do not put the key in code or commit it.
- If the improved version does not score higher, that is normal. Record the score, then change one thing at a time.
📚 Want to go deeper? Read Anthropic's Prompt Engineering Overview to understand “define success criteria first, then change the prompt”; then read the OpenAI Evals guide for a fuller evaluation workflow. For batch testing, explore
promptfoo/promptfoo. The complete resource list remains in Stage 2 Curated Projects, rather than being duplicated here.