|
|
I'm a researcher and engineer working on LLM agent reliability and program semantics. I study how agents use software and how program behavior changes across languages and execution environments.
My work connects evaluation design and systems implementation. I design experiments and build tools for tracing execution and analyzing programs, with the aim of making research results easier to inspect, explain, and reproduce.
I share some of my work here, alongside tools I build for research and everyday use. Rust and Python are my usual starting points; TypeScript and Swift come in for interfaces.
Research · Engineering · Projects · Collaboration
I build parsers, compiler tooling, and cross-language bindings to inspect and work with programs. I’m interested in where runtime behavior and language semantics affect what a transformation or an interface preserves.
Controlled environments Process tracing Benchmark harnesses
I build evaluation harnesses with controlled environments, process tracing, and inspectable execution records. They let me investigate failures and compare behavior across implementations and environments, connecting an experimental result to the system that produced it.
I build terminal interfaces, Tauri desktop apps, and native Swift/AppKit software. A shared Rust core lets me bring the same underlying behavior to different workflows, from a command line to a desktop interface.
I also work in established codebases, contributing to the compiler, packaging, fuzzing, and parsing tools that my own work depends on.
A few places to see how I work, or find something to build on:
- ⊢ Stepwise
My work on learning Python evaluation and natural deduction through individually checked steps.
- ⧉ Foch
Script analysis and merge tooling for Europa Universalis IV.
Some people play Europa Universalis IV. I also ended up writing tooling for its mod files. - ▦ Teaser
I’m building a macOS environment around complete project workspaces. - ⌘ ScriptMark
My tooling for grading programming assignments and inspecting the execution evidence behind a grade. - ◫ tree-sitter-paradox
A reusable grammar I maintain for Paradox scripts and their editor integrations.
An agent evaluation to untangle, a language tool to build, or an EU4 mod that refuses to merge?
I’m especially interested in collaborations that need experimental design and implementation together: developing an evaluation, investigating a failure, or building the tooling behind a research idea.
I'd like to hear about it. acturea@gmail.com



