test: prune redundant e2e matrix tests / test: 精简重复的端到端矩阵测试 - #3560
Conversation
合并端到端矩阵生成的重复测试,保留实验名称、上下文长度、并发列表及筛选行为的覆盖。
There was a problem hiding this comment.
Looks good, straightforward test consolidation with no production code changes. Reviewed: the exp-name/max-model-len merge (now checked per-row for both seq-len pairs across all concurrencies, still asserting exact expected values); the runner_node_filter-without-explicit-runner_type removal (test_runner_filter_keeps_each_commands_label_and_duplicate_policy uses the same runner_type=None default and asserts exact runner ordering/duplicates, parametrized over node_names including the empty case); and the multinode single-node-flag-skip removal (test_single_node_only_from_mixed uses a mixed single+multinode config and asserts no "prefill"-bearing entries survive, which is a stronger check than the deleted test's plain empty-result assertion). All claimed replacement tests exist in the file and genuinely exercise the same code paths.
Extended reasoning...
Single-file, test-only diff (+10/-56) in inferencex-e2e/infx/tests/matrix/test_generate_sweep_configs.py; no production code, no security-sensitive surface touched. Traced each of the three deleted tests (test_exp_name_format, test_max_model_len_calculation, test_runner_node_filter_without_runner_type, test_multinode_conc_as_list, test_single_node_flag_skips_multinode) against the fixtures and surviving tests to confirm the PR description's coverage-ownership table is accurate rather than aspirational. Found the merged/retained assertions actually reach the same code paths with equal or better precision (e.g., exact ordering assertions instead of a bare length check), so no coverage regression.
Description
Apply the OpenClaw test-audit skill to one focused matrix-generation batch in
inferencex-e2e/. Remove five redundant test functions and consolidate their contracts into existing behavioral tests.Only
inferencex-e2e/infx/tests/matrix/test_generate_sweep_configs.pychanges: +10/-56 lines, net -46 test lines; zero production or tooling changes.Coverage ownership
test_exp_name_formattest_sweep_expands_each_sequence_length_across_concurrenciesnow checks explicit experiment names for every 1k1k and 8k1k row.test_max_model_len_calculationtest_runner_node_filter_without_runner_typetest_runner_filter_keeps_each_commands_label_and_duplicate_policyexercises the full-sweep adapter with no explicit runner type and asserts exact runner ordering and duplicate behavior. Explicit-runner and no-match tests remain.test_multinode_conc_as_list[2150]assertion moves intotest_multinode_entry_structure, which uses the same arguments and fixtures.test_single_node_flag_skips_multinodetest_single_node_only_from_mixeduses the same multi-node fixture alongside valid single-node work, verifies nonempty output, and excludes every multi-node row.The audit covered the production adapters, shared expansion and row builders, callers, sibling tests, CI routing, and relevant history. No production seam can be removed by this batch. Schema, malformed-input, topology, parallelism, provenance, installed-package, and executable workflow-shell tests remain; other test owners are outside this PR.
Validation
Baseline: 441 matrix tests passed.
Final: 996 tests passed (436 matrix + 560 workflow tests):
Three temporary production mutations were each caught by retained tests: incorrect experiment name, context padding changed from 256 to 255, and bypassed topology filtering. All mutations were reverted before the final run; the production diff is empty.
Repository Ruff lint and format checks passed (
106 files already formatted);git diff --checkpassed. Ruff excludes the test directory by repository policy.Independent read-only review found no P0–P2 findings and separately passed all 436 matrix tests.
The Codex autoreview helper was unavailable in this session (
reviewer_unavailable/engine_failed); the independent review above is a fallback, not a successful Codex run.The first workflow run was launched outside the repository and had 20 checkout-root lookup failures. Running from the repository root passed all 560 workflow tests.
The full repository CPU suite has not been run locally; GitHub CI owns that broader gate. No GPU, Slurm, benchmark sweep, or evaluation dispatch was performed.
AI model disclosure
Type of Change
Checklist
中文
改动说明
按 OpenClaw 的 test-audit 技能,对
inferencex-e2e/中矩阵生成测试做一次小范围清理。删除五个重复测试函数,把仍有价值的断言合并到已有行为测试中。仅修改
inferencex-e2e/infx/tests/matrix/test_generate_sweep_configs.py:新增 10 行、删除 56 行,测试代码净减 46 行;生产代码和工具代码均无改动。[2150]的断言移入使用相同输入的条目测试。审查了生产适配器、共享展开逻辑、条目构造器、调用方、相邻测试、CI 路由和相关历史。没有可随本次测试清理删除的生产接口。配置校验、异常输入、拓扑、并行参数、来源追踪、安装包行为和实际执行工作流 shell 的测试均保留;其他模块不纳入本 PR。
验证
git diff --check通过。仓库 Ruff 配置排除了测试目录。reviewer_unavailable/engine_failed。独立审查是替代审查,不代表 Codex 审查成功。AI 模型使用说明
本次仅整理测试,不涉及文档、性能变更记录或 sweep 复用。本 PR 按请求保持未合并状态。