Skip to content

skillopt-sleep Codex harvest ingests its own headless replay sessions #286

Description

@kaluli123123

Summary

harvest_codex() accepts Codex sessions created by SkillOpt-Sleep itself during headless replay, judging, reflection, and tool-aware attempts. The other transcript harvesters explicitly reject engine-generated sessions, but the Codex source has no equivalent filter.

Reproduction

On main at 79124b37e9a6371e13b753f8bcd7adb1e493ade1, create a Codex archived session using the prompt shape emitted by CodexCliBackend.attempt_with_tools():

Complete the task. Apply the skill and memory rules EXACTLY, including any rule about searching before answering.

# Skill
...

# Memory
...

# Task
...

with a normal user_message and agent_message record. digest_codex_archived_session() returns a SessionDigest, and harvest_codex() includes it.

Expected behavior

Engine-generated Codex sessions should be excluded from harvested user evidence, just as Claude, Copilot CLI, and OpenCode replay/agent sessions are excluded.

Actual behavior

The generated prompt is treated as a real user task. On a later night it can be mined, replayed, and used to update the skill, creating self-training feedback and contaminating the task set.

Impact

Codex-backed runs can learn from their own evaluator/attempt instructions instead of user work, weakening task provenance and potentially reinforcing internal prompt text.

Environment

SkillOpt-Sleep main at 79124b37e9a6371e13b753f8bcd7adb1e493ade1; Python 3.10+; Codex CLI/archived sessions.

Activity

  1. kaluli123123 commented on Sep 20, 2026

    @kaluli123123
    Author

    I am working on a focused fix for this provenance bug. I will reuse the existing replay/session filters, add a narrow Codex-specific marker for the prompt shape emitted by SkillOpt's own Codex backend, and cover both exclusion and preservation of a real user prompt.

  2. kaluli123123 commented on Sep 20, 2026

    @kaluli123123
    Author

    Implemented in PR #287: #287

    Root cause: harvest_codex() did not distinguish SkillOpt-generated Codex replay prompts from user sessions, so internal headless runs could be harvested as user evidence. The fix reuses the agent-session filter and adds narrow Codex replay markers without applying the generic short-duration heuristic.

    Evidence:

    • Focused regression and existing harvest tests: 19 passed.
    • python3 -m compileall -q skillopt_sleep tests/test_harvest_codex_replay.py: passed.
    • git diff --check: passed.
    • Full python3 -m pytest -q remains unavailable because pytest is not installed in the environment.

    The PR is open; CI and maintainer review remain pending. No merge or Issue closure was performed.

  3. Yif-Yang commented on Sep 30, 2026

    @Yif-Yang
    Contributor

    #287 currently excludes the tool-enabled attempt shape, but synthetic transcripts built from the actual current attempt, judge, and reflect prompts are still harvested. The review also found that generic heading markers remove valid user sessions. Keep this issue open until those cases are covered; #282's new CLI session discovery makes the same provenance gap reachable in nested rollout files.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions