Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
22 changes: 21 additions & 1 deletion .github/workflows/ci.yml
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# Checks on every push and pull request.
#
# Split into three parallel jobs deliberately, so the wait is the slowest one rather than the
# Rust checks split into three parallel jobs deliberately, so the wait is the slowest one rather than the
# total: `check` is fast and catches most things, `board` catches the class of problem that only
# appears off the dev machine (cross-linking, glibc floors, unix-socket and permission
# semantics), and `coverage` needs its own instrumented build.
Expand Down Expand Up @@ -43,6 +43,17 @@ env:
RUSTFLAGS: -D warnings

jobs:
policy:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: astral-sh/setup-uv@v6
with:
version: "0.9.2"
python-version: "3.11"
- run: uv sync --locked --directory policy
- run: make -C policy lint test

check:
runs-on: ubuntu-latest
steps:
Expand All @@ -66,6 +77,15 @@ jobs:
- run: cargo fmt --all --check
- run: cargo clippy --workspace --all-targets
- run: cargo test --workspace
- uses: astral-sh/setup-uv@v6
with:
version: "0.9.2"
python-version: "3.11"
- name: Policy through real robotd and FakeIo
run: |
cargo build --locked -p robotd
uv sync --locked --directory policy
ROBOTD_BIN="$PWD/target/debug/robotd" uv run --directory policy pytest -q tests/test_robotd.py

# These are the pieces of this repo that run without being compiled — four on a robot,
# and `provision-board.sh` and `dev-push.sh` on a developer's machine — so nothing else
Expand Down
3 changes: 2 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,7 +18,8 @@ Text + body state/history → learned policy → action chunk → robotd → rob

Start with a language-conditioned motion policy, then add visual observations to explore
vision-language-action (VLA) models. The model learns the motion; `robotd` owns timed execution
and feedback. Training and model integration are ongoing work.
and feedback. A [replaceable Policy interface and Fake backend](policy/README.md) exercise this
execution path before training a model. Training and learned model integration are ongoing work.

## Origin

Expand Down
1 change: 1 addition & 0 deletions docs/README.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,7 @@
# Docs

Duckmind model execution: [Action Chunk API and timing contract](design/action-chunks.md).
Model boundary and runner: [Policy inference interface](design/policy-interface.md).

The [README](../README.md) is the front door — what a microduck is, and where to go. If you have
one in front of you and want to drive it, start at the [cheat sheet](robot/cheatsheet.md).
Expand Down
179 changes: 179 additions & 0 deletions docs/design/policy-interface.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,179 @@
# Policy inference interface

The model boundary is `Policy.infer(PredictRequest) → ActionChunk`, implemented in
[`policy/`](../../policy/README.md). A rule, local neural network, or remote model adapter can
implement it. Inference does not acquire a robot or write motors; the runner owns a robotd
Action Session and submits predictions to the existing executor.

```text
text + latest body observation + history
↓
Policy.infer / HTTP service FakePolicy today; learned adapter later
↓
decoded joint ActionChunk
↓
Runner on robotd's machine discard expired prefix; submit remaining timeline
↓ Unix JSON-RPC
robotd → Safety → RobotIo → body
↑
robot.state
```

This separates the **model contract** from the **execution contract**. The latter, including
control ownership, admission, replacement, stops and hold behavior, belongs to
[Action Chunks](action-chunks.md). No Rust IPC or control-loop behavior changes in this layer.

## Model contract: microduck-joints-v1

The Pydantic types in [`schema.py`](../../policy/src/duckmind_policy/schema.py) are the source
of truth; FastAPI publishes their JSON Schema through `/openapi.json`. This is a Duckmind
body-specific contract, not an industry standard and not a promise of cross-embodiment control.

`POST /v1/policy/predict` accepts:

| Field | Meaning |
|---|---|
| `schema_version` | `microduck-joints-v1`; defaults to this value, other versions reject. |
| `task` | Free-form nonempty text. The API has no command enum. |
| `observation.t_ns` | Positive robotd monotonic timestamp in nanoseconds, from `robot.state`. |
| `observation.positions` | 15 measured absolute joint angles, radians, in robotd `JOINT_NAMES` order. |
| `observation.velocities` | Optional 15 measured velocities, rad/s; absent means unavailable. |
| `observation.imu` | Optional `gyro` (trunk rad/s), `quat` (trunk→world, scalar-first wxyz). |
| `observation.images` | Optional map of camera name to `{t_ns, jpeg_base64}`; image capture time uses the same robot clock and cannot be newer than the observation. |
| `history` | Up to eight earlier observations in strictly increasing timestamp order; excludes the current sample. |

History is explicit and each inference request is independent: there is no server-side
conversation or hidden per-robot temporal state. A model needing a context window builds it
from these samples; a model needing a different history contract requires a versioned change.
The stock runner sends recent body states only, with IMU/velocities when available. Camera
capture, synchronization, and JPEG conversion are not implemented by this runner yet. An
image-aware collector should use the existing WebRTC consumer path, preserve capture times,
and form the same `PredictRequest`; do not invent a separate camera transport.

The response contains:

| Field | Meaning |
|---|---|
| `schema_version` | `microduck-joints-v1`. |
| `observation_t_ns` | Exact timestamp from the request; ties the prediction to its observation. |
| `step_ns` | Exactly 20,000,000 ns (50 Hz). |
| `positions` | H × 15 finite absolute-radian targets, 1 ≤ H ≤ 100. |

Frame `i` belongs to `[observation_t_ns + i*step_ns, observation_t_ns + (i+1)*step_ns)`.
Inference latency consumes part of this horizon. The response contains no session ID or
execution sequence: those belong to the runner, not the model. A response predicts motion;
it does not report task success.

A model adapter owns preprocessing, image decoding/resizing, state normalization, inference,
action denormalization and joint mapping. It must return this decoded action space. Upstream
Microduck's 14 normalized locomotion outputs are **not** directly compatible: the adapter must
apply its model's offsets/scales and supply all 15 targets, including the mouth. The ±π wire
bound is only a coarse numeric envelope, not collision checking or a balance guarantee.

`GET /v1/policy` exposes joint order, timestep and action-space metadata. The runner also checks
robotd's acquisition response against its expected joint order/timestep before inference.

## Replace a backend

The Python protocol has one method. No particular model architecture, trainer or GPU is required.

```python
from duckmind_policy.policy import FakePolicy, Policy
from duckmind_policy.schema import ActionChunk, PredictRequest
from duckmind_policy.service import HttpPolicy, create_app
import uvicorn

# Local implementation, useful before training exists:
policy: Policy = FakePolicy()

# The HTTP client implements the very same interface:
remote: Policy = HttpPolicy("http://127.0.0.1:8081")

# A future adapter implements this signature and performs its model-specific conversions:
# def infer(self, request: PredictRequest) -> ActionChunk: ...
uvicorn.run(create_app(policy), host="127.0.0.1", port=8081)
```

The built-in FakePolicy supports `向左看`/`往左看`/`look left`, right equivalents, and
`保持姿势`/`hold`. It copies observed non-head joints and approaches an absolute head-yaw target
at at most 0.5 rad/s in its predictions. Unsupported tasks fail explicitly; these aliases are
only the fake backend's test vocabulary. A learned implementation can accept any task it has
learned without changing the API.

The server serializes inference and rejects overlap with HTTP 503 rather than queueing old
observations. Invalid requests/unsupported tasks return 422, backend failures return 500.
The HTTP client has one pooled connection, configured I/O timeout, and no automatic retries.
A timed-out server inference may finish, but its result cannot submit itself to robotd. The
service binds to loopback by default; remote access/authentication deployment is outside this
prototype. The runner and robotd remain on the same machine; a GPU policy service can be remote.

## Runner timing and lifecycle

The runner acquires one Action Session, drains `robot.state` on a separate thread, and runs one
inference request at a time. It keeps only the latest feedback and rejects feedback older than
250 ms. It uses robot timestamps plus locally elapsed receive time; wall clocks and the model
server's clock do not determine execution time.

After inference, it verifies the session is still active, removes expired frames plus two ticks
of admission lead, and submits the remaining targets with their **original timestamps**.
It never rebases old predictions to "now". A completely expired response fails. Replanning occurs
at most 200 ms after admission (earlier for a short remaining horizon); robotd retains the prior
prefix until replacement starts. Models must produce a horizon long enough for inference,
admission and the next prediction. A 1–2-frame response is schema-valid but too short for this
runner's two-tick lead; slow inference can exhaust even a longer horizon. Admission failures stop
the run rather than silently skipping, retrying or acquiring a new session.

Normal duration completion, inference failure, Ctrl-C and lost feedback all leave through a
`finally` block that ends this runner's session. An external stop invalidates it; a late result
is discarded or rejected by robotd. Cleanup never issues a global stop against a newer owner.
If the process dies or IPC is unavailable, robotd's own bounded timeline/expiry remains the
fallback. Session end/expiry holds the last pose; it is not an active balance or recovery policy.

## Run and verify

Build robotd using the repository's Rust prerequisites, then start an isolated fake instance:

```sh
cargo build -p robotd -p robotctl
printf '[audio]\nenabled=false\n[chorale]\naccept=false\n' > /tmp/duckmind-policy.toml
./target/debug/robotd --fake --no-policy --socket /tmp/duckmind-policy.sock \
--params /tmp/duckmind-policy.toml
```

In another terminal, initialize it and allow the home transition to finish:

```sh
./target/debug/robotctl --robot-socket /tmp/duckmind-policy.sock robot init
cd policy
uv sync --locked
uv run duckmind-policy serve
```

Then start inference/execution:

```sh
cd policy
uv run duckmind-policy run --socket /tmp/duckmind-policy.sock --task '向左看' --duration 5
```

Unit/HTTP tests run with `make test`. Setting `ROBOTD_BIN` also runs real-daemon tests that
launch their own fake robotd, initialize it, serve HTTP, submit/replan chunks, assert measured
joint feedback, and verify a stop during inference rejects the late result. Without that variable
those tests report skips, not success. CI sets it after building robotd.

This validates the software path using FakeIo. A frozen-leg head-turn rule is not a MuJoCo
walking/balancing policy; neither physics performance, camera-driven VLA behavior nor a physical
robot is validated here. The existing robotd Action Chunk API remains restricted to fake/sim.

## Precedents

[OpenPI BasePolicy](https://github.com/Physical-Intelligence/openpi/blob/main/packages/openpi-client/src/openpi_client/base_policy.py)
provides an interchangeable `infer(obs)` boundary and its WebSocket client implements that same
abstraction. [OpenPI remote inference](https://github.com/Physical-Intelligence/openpi/blob/main/docs/remote_inference.md)
separates the model server from robot execution.
[LeRobot PreTrainedPolicy](https://github.com/huggingface/lerobot/blob/main/src/lerobot/policies/pretrained.py)
exposes `predict_action_chunk`, and its
[async server tests](https://github.com/huggingface/lerobot/blob/main/tests/async_inference/test_policy_server.py)
use a mock policy. Duckmind borrows the separation and replaceability, not their wire protocols
or dependencies. Continuous execution during synchronous remote inference is supported; RTC,
parallel inference requests and model training are not implemented by this package.
5 changes: 5 additions & 0 deletions policy/.gitignore
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
.venv/
__pycache__/
.pytest_cache/
.mypy_cache/
.ruff_cache/
10 changes: 10 additions & 0 deletions policy/Makefile
Original file line number Diff line number Diff line change
@@ -0,0 +1,10 @@
.PHONY: fix lint test
fix:
uv run ruff format .
uv run ruff check --fix .
lint:
uv run ruff check .
uv run ruff format --check .
uv run mypy src
test:
uv run pytest -q
30 changes: 30 additions & 0 deletions policy/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,30 @@
# Duckmind Policy

A replaceable `Policy.infer(request) → ActionChunk` boundary. The same request/response can
run locally or over HTTP; the runner connects predictions to the existing robotd executor.
The included FakePolicy exercises plumbing with head turns and hold. It is not a trained VLA.

From this directory:

```sh
uv sync --locked
uv run duckmind-policy serve
# Another terminal, after initializing robotd --fake or the simulated robot:
uv run duckmind-policy run --socket /tmp/robot.sock --task '向左看' --duration 5
```

`--url` selects another compatible policy service; `--timeout` bounds HTTP I/O (default 0.8 s).
The model service exposes `POST /v1/policy/predict`, `GET /v1/policy`, and OpenAPI at `/docs`.
To serve a learned adapter, implement `Policy.infer` and pass it to `create_app(adapter)`.
No runner or robotd changes are required if the adapter meets the contract.

See [the Policy interface design](../docs/design/policy-interface.md) for schemas, timing,
backend replacement, simulation setup, and current limits.

```sh
make fix
make lint
make test
# Run the real HTTP → robotd → FakeIo integration tests too:
ROBOTD_BIN="$(pwd)/../target/debug/robotd" make test
```
29 changes: 29 additions & 0 deletions policy/pyproject.toml
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
[project]
name = "duckmind-policy"
version = "0.1.0"
description = "Replaceable action-chunk policies for Duckmind"
requires-python = ">=3.11"
dependencies = ["fastapi>=0.115,<1", "httpx>=0.28,<1", "pydantic>=2.10,<3", "uvicorn>=0.34,<1"]

[project.scripts]
duckmind-policy = "duckmind_policy.cli:main"

[build-system]
requires = ["hatchling"]
build-backend = "hatchling.build"

[dependency-groups]
dev = ["pytest>=8,<10", "ruff>=0.11,<1", "mypy>=1.15,<2"]

[tool.ruff]
target-version = "py311"
line-length = 100

[tool.ruff.lint]
select = ["E", "F", "I", "UP", "B"]

[tool.mypy]
strict = true

[tool.pytest.ini_options]
testpaths = ["tests"]
1 change: 1 addition & 0 deletions policy/src/duckmind_policy/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
"""Model-independent policy inference and the robotd execution bridge."""
37 changes: 37 additions & 0 deletions policy/src/duckmind_policy/cli.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,37 @@
"""Local serving and a bounded simulation runner."""

import argparse
import logging

import uvicorn

from duckmind_policy.policy import FakePolicy
from duckmind_policy.robot import RobotClient
from duckmind_policy.runner import run
from duckmind_policy.service import HttpPolicy, create_app


def main() -> None:
parser = argparse.ArgumentParser(description=__doc__)
commands = parser.add_subparsers(dest="command", required=True)
serve = commands.add_parser("serve", help="serve the built-in FakePolicy")
serve.add_argument("--host", default="127.0.0.1")
serve.add_argument("--port", type=int, default=8081)
runner = commands.add_parser("run", help="run an initialized fake/sim robot through HTTP")
runner.add_argument("--socket", required=True)
runner.add_argument("--url", default="http://127.0.0.1:8081")
runner.add_argument("--task", required=True)
runner.add_argument("--duration", type=float, default=5.0)
runner.add_argument("--timeout", type=float, default=0.8, help="HTTP timeout in seconds")
args = parser.parse_args()
logging.basicConfig(level=logging.INFO)
if args.command == "serve":
uvicorn.run(create_app(FakePolicy()), host=args.host, port=args.port)
else:
policy = HttpPolicy(args.url, timeout=args.timeout)
try:
run(RobotClient(args.socket), policy, args.task, args.duration)
except KeyboardInterrupt:
pass # run's finally block releases its action session.
finally:
policy.close()
Loading
Loading