AgentEvalTool/docs/adr/0008-deepen-intelligent-eval-seams.md
sinohqb 37da87c3b0
Some checks failed
CI / test (push) Failing after 4m30s
refactor(intelligent-eval): openclaw_client token from settings (S5) + ADR-0008
P4 boundary (issue #10):
- S5: openclaw_client reads gateway_token and container_name from
  Settings (AGENTEVAL_OPENCLAW_GATEWAY_TOKEN / CONTAINER_NAME), with
  backwards-compatible defaults. Test injection still works via __init__.
- ADR-0008 documents the full deepening: S1 (domain convergence), S4
  (stuck-task settlement), S2 (router → service), S3 deferral rationale,
  S5 (token config), S6 (frontend polling deferred).

S6 (CronPoolMonitor unified polling) deferred to independent issue.
No behaviour change — 873 passed + 5 xfailed unchanged.
2026-08-13 14:20:49 +08:00

29 lines
3.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# ADR-0008: Deepen intelligent-eval Scheduling, Task Settlement, and Router Seams
**Status**: Accepted
**Date**: 2026-08-13
**Deciders**: AgentEval Team
## Context
v1.1.0 introduced the `intelligent_eval/` module (Cron Pool architecture, 13 files, 2883 lines). The module was built as feature work (ticket 0110) and embedded several architectural seams consistent with rapid delivery:
- **S1** — time-slot / deficit / severity domain knowledge was duplicated between `task_queue.py` and `decision.py` (two independent implementations of "8-10h" parsing, slot counting, and priority calculus).
- **S4** — `cron_pool.handle_stuck_cron` performed a runtime `from ... import complete_task` to settle stuck-cron tasks, coupling the two modules.
- **S2** — router handlers (`intelligent_evals.py` and `openclaw_cron_pool.py`) directly executed ORM writes (decision-logs, heartbeat, task-queue eval loading, scale direction decision).
- **S3** — `decision.py` and `task_queue.py` directly accepted `Session` and queried `IntelligentEvalSessionDB` (DB coupling). The internal functions were already testable via `db_session` fixture.
## Decision
- **S1 (converge scheduling domain)**: Created `agenteval.intelligent_eval.domain.py` as the single source of truth for `parse_time_slot`, `is_slot_due`, `get_current_slot`, `count_sessions_in_slot`, `count_total_sessions`, `calculate_session_deficit`, `calculate_priority`, `get_attention_reason`, and `has_high_severity_issues`. `task_queue` and `decision` now delegate to `domain` via thin wrappers that preserve the original `_` function signatures, keeping existing tests compatible.
- **S4 (converge stuck-task settlement)**: Moved task settlement logic into `task_queue.requeue_stuck_task(eval_id, cron_id, session)`; `cron_pool.handle_stuck_cron` calls it, removing the runtime import.
- **S2 (router logic down to service)**: Extracted `decision_logs.create_decision_log` / `list_decision_logs`; added `cron_pool.heartbeat` and `cron_pool.scale_to`; added `task_queue.get_next_task_with_eval`. Router handlers now only validate HTTP inputs and translate `LookupError` → 404.
- **S3 (DB decoupling)**: DELIBERATELY DEFERRED. The `domain` functions already accept `Session` and the test suite already exercises them through `db_session`. Pulling a full repository/read-model abstraction layer at this point would add indirection without a second adapter that benefits from it. The existing 873-test suite serves as the behavioural lock.
- **S5 (openclaw_client token configuration)**: `OpenClawClient.__init__` now reads `gateway_token` and `container_name` from `Settings` (`AGENTEVAL_OPENCLAW_GATEWAY_TOKEN` / `AGENTEVAL_OPENCLAW_CONTAINER_NAME`) when not explicitly passed. The previous hardcoded values remain as defaults.
- **S6 (frontend CronPoolMonitor polling)**: Deferred to a separate issue (the `it.fails` guard in `CronPoolMonitor.test.tsx` documents the gap).
## Consequences
- **Positive**: All three deep seams (S1, S4, S2) converged with zero observable behaviour change (873 passed + 5 xfailed unchanged). The module now has a single scheduling domain, stuck-task settlement is a first-class operation, and router handlers are thin HTTP translators.
- **Negative**: The S3 deferral means `domain` functions still accept a `Session` parameter; pure logic cannot be unit-tested without a DB. This is acceptable until a second adapter (e.g., a different storage backend) demands the abstraction.
- **Open**: S6 (frontend polling unification) is tracked as an independent issue; the `it.fails` guard in the test suite will alert when the gap is fixed.