sinohqb
71d38ebef6
fix(intelligent-eval): settle stale tasks of finished evals
...
CI / test (push) Successful in 3m57s
任务队列出现'评估已 completed 却有待认领/执行中任务'的残留:评估离开
executing 后,其 pending/assigned 任务无人清理(requeue_stale_assigned_tasks
只处理 executing 评估的 assigned,评估结束被跳过)。
- 新增 task_queue.settle_tasks_for_finished_evals:评估非 executing 时,
其 pending/assigned 任务回收为 completed;executing 评估的任务保留
- scan loop 每 60s 在 requeue+scan 后调用清理
测试:+1(已结束评估的 pending/assigned 回收、executing 保留、幂等),901 passed
2026-08-18 16:04:46 +08:00
sinohqb
b4f9c887f4
feat(intelligent-eval): task queue monitor (方案③可视化)
...
CI / test (push) Successful in 4m1s
方案③的定时触发(scan loop 每 60s 入队 + 触发 worker)此前只有 Worker
消费端 API,无可查看的列表。新增:
- GET /api/intelligent-evals/tasks:任务明细(含评估名/状态)+ 状态分布统计
(注册在 /{eval_id} 之前避免被捕获为 eval_id="tasks")
- 前端 TaskQueueMonitor 组件 + 智能评估页任务队列入口:5s 轮询
(usePolling),状态卡 + 状态筛选 + 明细表
测试:+3(列表/筛选/不被 {eval_id} 遮蔽),892 passed,tsc 通过
2026-08-17 13:57:00 +08:00
sinohqb
71c7cd3d39
fix(intelligent-eval): requeue stale assigned tasks (worker crash recovery)
...
CI / test (push) Successful in 3m45s
方案③ worker 由平台触发 openclaw agent(cron=manual-run,非真实 cron),
fault_tolerance 的 stuck 检测不适用——agent 中断/失败时任务永久卡 assigned,
scan 只查 pending 不再入队(死锁)。
requeue_stale_assigned_tasks:assigned 超过 10 分钟且评估仍 executing 的
任务重置为 pending(清空认领),平台 scan 循环随后重新触发 worker 重试。
接入 scan loop,每轮先 requeue 再 scan。
2026-08-17 03:30:42 +08:00
sinohqb
38e3817433
fix(intelligent-eval): atomic CAS in assign_task and complete_task (resolves §6.1)
...
CI / test (push) Has been cancelled
Replace read-check-write in task_queue.assign_task with UPDATE...WHERE
status='pending' and decide on rowcount so two concurrent workers
cannot both claim the same task. Also harden complete_task with the
same CAS pattern (status='assigned') so a worker + stuck-handler
double-complete leaves the DB in one state.
The xfail guard in test_worker_task_resilience now passes (4/4).
2026-08-14 14:52:34 +08:00
sinohqb
3852c6f87d
refactor(intelligent-eval): router logic down to service layer (P3, S2)
...
CI / test (push) Failing after 4m16s
P3 deepening (issue #9 ): remove direct ORM from router handlers.
- decision_logs.py (new): create_decision_log / list_decision_logs service
- cron_pool.heartbeat: encapsulate heartbeat cron lookup + update + commit
- cron_pool.scale_to: encapsulate scale direction decision (if/elif/else)
- task_queue.get_next_task_with_eval: encapsulate eval-loading + dict-building
- Router endpoints now delegate to services, only handling HTTP-level
validation (status codes, 404 translation via LookupError).
No observable behaviour change — 873 passed + 5 xfailed unchanged.
T3 router ORM contract guards (5/5) continue to pass.
2026-08-13 13:57:34 +08:00
sinohqb
975ed7a6ff
refactor(intelligent-eval): converge scheduling domain (S1) + stuck-task settlement (S4)
...
CI / test (push) Failing after 4m21s
P1 deepening (issue #7 ):
S1: is the single source of truth for time-slot parsing,
slot-due checks, session deficit, priority, attention reason, and
high-severity detection. and now delegate
their internal helpers to while keeping the same signatures
(tests continue to pass via the thin wrappers).
S4: encapsulates the stuck-cron
task settlement (fail current task + enqueue retry).
calls it instead of the previous runtime import of .
No observable behaviour change — 873 passed + 5 xfailed unchanged.
2026-08-13 10:04:24 +08:00
sinohqb
4e9145db46
fix(lint): resolve ruff lint errors
...
CI / test (push) Failing after 32s
- Remove unused imports (json, datetime, timedelta, Any, Optional, IntelligentEval)
- Remove unused variable (estimated_sessions)
- Remove whitespace from blank line
- Organize import blocks
All ruff checks passing.
2026-08-12 11:13:19 +08:00
sinohqb
1aa453ef0a
feat(intelligent-eval): add cron pool data model and task queue API
...
Implement Ticket 01 of intelligent eval cron pool architecture (ADR-0007):
- Add 4 new tables: task_queue, cron_pool, config_snapshots, decision_logs
- Implement task enqueueing logic with priority calculation
- Implement task assignment and completion APIs
- Add unit tests (9) and integration tests (7)
- Update CONTEXT.md with new vocabulary
- Add ADR-0007 documenting cron pool architecture decision
All 760 tests passing.
2026-08-12 02:13:21 +08:00