Skip to content

Commit 43412e7

Browse files
Stephen Gruppettaclaude
andcommitted
Rename folder to llm-evaluation and link the tutorial in the README
Match the house convention of materials folder = article slug, which the Trello card gives as "llm-evaluation" (keyphrase "llm evaluation"). Also add the tutorial link to the README, which was missing, and note the follow_along/completed split plus how to run in replay and live mode. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
1 parent 7b89e6f commit 43412e7

35 files changed

Lines changed: 26 additions & 8 deletions

‎build-llm-evaluation-harness-python/README.md‎

Lines changed: 0 additions & 8 deletions
This file was deleted.

build-llm-evaluation-harness-python/.gitignore renamed to llm-evaluation/.gitignore

File renamed without changes.

‎llm-evaluation/README.md‎

Lines changed: 26 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,26 @@
1+
# LLM Evaluation in Python: Build an Eval Harness From Scratch
2+
3+
This folder contains the source code for [LLM Evaluation in Python: Build an Eval Harness From Scratch](https://realpython.com/llm-evaluation/).
4+
5+
- `follow_along/` contains the prompts, replay corpus, and setup files that you should start with.
6+
- `completed/` contains the finished evaluation harness.
7+
8+
## Setup
9+
10+
Move into `follow_along/`, then create your virtual environment and install the dependencies:
11+
12+
```console
13+
$ uv sync
14+
```
15+
16+
## Usage
17+
18+
Replay mode is the default and needs no API key, since every command reads recorded generations and judgments from `data/replay/`:
19+
20+
```console
21+
$ uv run evals.py run --prompt prompts/support_v1.txt
22+
```
23+
24+
Live mode calls the OpenAI API through the same `ModelClient` protocol. Copy `.env.example` to `.env`, add your key, then pass `--live`.
25+
26+
See the tutorial for the full walkthrough.

build-llm-evaluation-harness-python/completed/.env.example renamed to llm-evaluation/completed/.env.example

File renamed without changes.

build-llm-evaluation-harness-python/completed/data/judge_calibration.jsonl renamed to llm-evaluation/completed/data/judge_calibration.jsonl

File renamed without changes.

build-llm-evaluation-harness-python/completed/data/replay/responses.json renamed to llm-evaluation/completed/data/replay/responses.json

File renamed without changes.

build-llm-evaluation-harness-python/completed/data/support_eval.jsonl renamed to llm-evaluation/completed/data/support_eval.jsonl

File renamed without changes.

build-llm-evaluation-harness-python/completed/deepeval_example.py renamed to llm-evaluation/completed/deepeval_example.py

File renamed without changes.

build-llm-evaluation-harness-python/completed/evals.py renamed to llm-evaluation/completed/evals.py

File renamed without changes.

build-llm-evaluation-harness-python/completed/llm_eval/__init__.py renamed to llm-evaluation/completed/llm_eval/__init__.py

File renamed without changes.

0 commit comments

Comments
 (0)