Shaurya Singh · Seattle, WA · cs + data science at UW, class of '27 · shauryasingh.com
i find what hurts, then build at it. so far: proteins, brains, quantum noise, classrooms, street safety, buildings, an underwater robot, and the code that grades things.
$ grep -rl "same field" ~/built/
no matches
| since 2026-09-23 | ||
|---|---|---|
| issues filed | 238 | in 229 repositories, 15 areas |
| fixed by a merged pull request | 20 | 17 by other people, 3 by me |
| median hours to a fix | 109.5 | filed to merged, over 17 fixes by others |
| pull requests to other people's repositories | 33 | 3 merged, 28 open |
the habit: i read the code that decides whether an answer is right (eval harnesses, classroom autograders, forecasting and fairness metrics) and file what i find. since 2026-09-23: 238 issues in 229 repositories, 20 fixed by a merged pull request so far, 17 of those by other people and 3 by me.
one example, small enough to read: scoreprobe on a two-line grader, before and after a one-line fix.
three more panels and the one-defect table
eight of the fixes share one defect: a missing, empty or undefined result counted as a score.
| report, fixed by | what was wrong |
|---|---|
| otter-grader#1020 shaurya416 #1021 | grading produced no results and the log still said all tests passed |
| AReaL#1757 MohammadHijjawi97 #1758 | a reply with no answer scored 0.9 against a "None" placeholder gold |
| evalscope#1783 Yunnglin #1785 | a blank answer against a blank reference scored full credit |
| mlxtend#1202 MohammadHijjawi97 #1203 | mcnemar returned p = 0.0 by default when there were no discordant pairs |
| utilsforecast#285 vgvr0 #286 | smape scored a missing forecast as a perfect 0 |
| trl#7364 adithya-s-k #7467 | OpenReward's default GRPO reward could not tell never scored from scored 0.0 |
| Curator#2443 1fanwang #2444 | an empty document scored 1.0 on the repetition filter and was kept |
| darts#3221 adigulalkari #3223 | grid search returned a NaN-scored combination as the best when it came first in the grid |
19 merged pull requests closed 20 of the issues
one line per world i've wandered into.
- proteins · UW's DAIS lab: preprocessing, training runs and renders for DeepTracer, cryo-EM maps to 3D protein structures. no byline, and that's fine.
- brains · four essays, taught to myself in public, optogenetics to connectomics (2022–23). now trigalign, for EEG recordings that dropped a trigger or two.
- quantum · shotfloor: the score a perfect quantum computer gets from shot noise alone. 8 qubits, uniform target, 1,024 shots: 0.92 hellinger fidelity, not 1.
- classrooms · co-founded SkillTern, its CTO (2022–23): computer science for 350+ students, built to run without me.
- street safety · TurtleShell: K-means over 10,000+ LAPD crime datapoints and an iOS SOS app. Microsoft for Startups, then shut down on purpose (oct 2023).
- buildings · a cabin designer that has to pass the Seattle residential code: rule engine, NSGA-II over cost, space and comfort, IFC export. runs locally, so no link.
- earlier · H2OAquatics, an autonomous underwater vehicle aimed at ocean acidification; a phishing policy paper, runner-up at WatGov (Waterloo, 2023).
- off the clock · music. made often, released never.
small python tools, each about one thing that quietly goes wrong. none is public yet, so no links: each name becomes one the day its repository is.
| tests and repos | tolwatch | pytest plugin: flags allclose, isclose and approx checks that can't fail, or are far looser than they need to be |
| ziptrace | names the zip() and map() calls that silently dropped data to a length mismatch |
|
| unadded | finds the untracked or gitignored files a passing test run quietly depends on, before CI does | |
| whymodified | git status says modified and nothing changed: says why, and the one command that clears it |
|
| science data | shotfloor | the score a flawless quantum computer gets at a given shot count, and the shots it needs to reach a target |
| trigalign | pairs a stimulus log with an EEG trigger list when triggers dropped, doubled, or the clocks drifted | |
| grading code | scoreprobe | feeds a grader the inputs known to slip through (a blank answer, "not Paris" against "Paris", 4200 against 42) and reports which earned credit |
| normladder | is "1,000" the same answer as "1000"? 8 of 14 answer-matching rules from lm-eval, HELM and SQuAD say yes; it shows which, and why | |
| mcqlint | lints multiple-choice eval sets: keys that name no option, duplicate options, leaked answers, position bias | |
| tabverify | recomputes the averages, deltas and bold in results tables at the printed precision, and flags what no rounding explains |
method
- counts are
total_countfrom GitHub search at 2026-10-08 20:50 utc:author:shaurya416 is:public is:issue(238: 213 open, 25 closed) andauthor:shaurya416 is:public is:pr(33: 3 merged, 28 open, 2 closed unmerged; 33 to other people's repositories), excluding the owner's own organisations. - fixed means the issue is closed and a pull request in its repository that references it was merged before it closed; an issue closed as not planned never counts. all 25 closed issues were traced: 20 fixed, 5 closed without a merged pull request. the 16 pull requests by others come from 14 people.
- an open issue whose fix did not close it is not counted, so fixed is a lower bound.
- hours run from the issue being filed to the pull request merging.
- areas: each of the 229 repositories is put in one of 15 areas by hand. that, the four kinds of code the side-quest sentence names, the sentence above the defect table and its one line per row are the hand-written inputs to the side quest.
- the scoreprobe panel is a captured run of scoreprobe on the code it shows. its first public release is pending; with its source,
python3 build.py capture-probe --scoreprobe-src PATHreruns the panel. - days are utc. a day with nothing is drawn as a mark, not left blank.
- regenerate:
python3 build.py fetchreads the api (it needsGITHUB_TOKEN), thenpython3 build.pyrenders. same input, same bytes.
side-quest numbers: github api, as of 2026-10-08 20:50 utc.
