• Joined on 2026-06-26
solahqb commented on issue solahqb/AgentEvalTool#1 2026-08-11 08:44:51 +00:00
Refactor: Deepen evaluation architecture modules

已完成并验收。核心架构深化由提交 c896ab3 完成,生命周期一致性修正由 6248568 完成;最终随 v1.0.0(0326ec5)通过 744 项后端测试、15 项前端测试、Ruff…

solahqb pushed tag v1.0.0 to solahqb/AgentEvalTool 2026-08-11 08:04:02 +00:00
solahqb pushed to main at solahqb/AgentEvalTool 2026-08-11 08:03:43 +00:00
0326ec5d03 fix(deploy): serialize production frontend build
941df4c8c7 release: normalize project version to v1.0.0
1782b245bf refactor(architecture): deepen campaign runtime modules
Compare 3 commits »
solahqb commented on issue solahqb/AgentEvalTool#2 2026-08-10 08:16:01 +00:00
Refactor: Deepen Campaign Intelligence, Read Model, and Runtime

Phase 0 基线完成:后端 Campaign/分析/周期对比/报告/runtime 针对性测试 251 passed;前端新增 useCampaignReport 行为基线 3 项后全套 12 passed;生产实现暂未修…

solahqb opened issue solahqb/AgentEvalTool#2 2026-08-10 08:10:36 +00:00
Refactor: Deepen Campaign Intelligence, Read Model, and Runtime
solahqb released AgentEvalTool v0.8.0 at solahqb/AgentEvalTool 2026-08-09 05:57:55 +00:00
solahqb pushed tag v0.8.0 to solahqb/AgentEvalTool 2026-08-09 05:44:36 +00:00
solahqb pushed to main at solahqb/AgentEvalTool 2026-08-09 05:44:17 +00:00
10a089e740 docs(deploy): record production access topology
864ae2b7fe fix(review): address release correctness findings
4a0709b456 fix(frontend): synchronize production lockfile
855455c0b4 feat(deploy): add production release workflow
62485684ca fix(architecture): enforce lifecycle consistency
Compare 6 commits »
solahqb opened issue solahqb/AgentEvalTool#1 2026-08-06 07:09:57 +00:00
Refactor: Deepen evaluation architecture modules
solahqb pushed to main at solahqb/AgentEvalTool 2026-08-05 18:02:31 +00:00
f26c34a340 docs(ui-consistency): archive completed refactor, record drawer evolution in ADR-0005
804880f1e3 fix(intelligent-eval): add top padding in drawer bodies
1fbd2bceed docs(ui-consistency): check off all ticket acceptance items after t480 walkthrough
9f25728c2f docs(agents): detail views may use right-side Drawer per ADR-0005
01f451c155 feat(intelligent-eval): align detail/report drawer layouts with other modules
Compare 17 commits »
solahqb pushed to main at solahqb/AgentEvalTool 2026-08-05 06:01:44 +00:00
6cc2efafb6 feat(frontend): detail auto-refresh, session list, pinned AI-assistant tab
9cdbc41808 docs(intelligent-eval): commit v1.0 spec, tickets, and post-v1.0 improvements
Compare 2 commits »
solahqb pushed to main at solahqb/AgentEvalTool 2026-08-04 20:07:52 +00:00
e1491dfc97 feat(intelligent-eval): OpenClaw planner/evaluator/analyst skills (ticket 08)
solahqb pushed to main at solahqb/AgentEvalTool 2026-08-04 19:49:45 +00:00
cbee749da2 feat(intelligent-eval): frontend list/detail/report pages (tickets 05-07)
solahqb pushed to main at solahqb/AgentEvalTool 2026-08-04 19:34:11 +00:00
da6ccc265d feat(frontend): restructure admin menu into two-level groups and add intelligent-eval entry
1317552701 feat(intelligent-eval): add backend for OpenClaw-driven intelligent evaluation (tickets 01-04)
Compare 2 commits »
solahqb pushed to main at solahqb/AgentEvalTool 2026-08-04 05:40:21 +00:00
f93d2f224f fix(frontend): auto-fetch report data when campaignId changes
solahqb pushed to main at solahqb/AgentEvalTool 2026-08-04 05:29:39 +00:00
160332665e refactor(exploration): absorb settlement.py into ExplorationSessionRepository
0aa3ef81c5 refactor(repository): extract AsyncJobRepository base class for analysis/comparison
cbfdf86b36 refactor(frontend): extract useCampaignReport hook from Campaigns.tsx
c24998c762 refactor(metrics): extract dashboard aggregation to compute_dashboard
42be31dd1f feat(report): add load_campaign_view as unified campaign read model
Compare 8 commits »
solahqb pushed to main at solahqb/AgentEvalTool 2026-08-03 19:47:47 +00:00
df76edcf55 refactor(exploration): move ledger and state machine into domain modules
38849d46f1 refactor(repository): narrow atomic updates for patrol/cancel/scheduler writes
f8d8450b1e refactor(report): unify campaign report loading behind one read model
Compare 3 commits »
solahqb pushed to main at solahqb/AgentEvalTool 2026-08-03 18:55:51 +00:00
a665b496b0 chore(v0.9): wrap over-length lines and record spec rulings
b26c432f3a refactor(frontend): reuse ChatBubble for exploration drill-down
ef4c094082 refactor(exploration): share one fetch+aggregate helper across outlets
936640fb36 fix(exploration): include all judge findings instead of poor-only
7f68afd765 docs(v0.9): mark ticket 06 done after Playwright browser verification
Compare 18 commits »
solahqb pushed to main at solahqb/AgentEvalTool 2026-08-03 10:24:16 +00:00
3e485bfbe6 docs(v0.9): mark ticket 03 patrol API done
6340ec503c feat(exploration): stateless patrol API with watermark increments
53afb9b5d1 feat(campaigns): exploration seed set and budget config per campaign
a351f65550 feat(exploration): session lifecycle endpoints with platform hard guardrails
c35159ffcb docs(v2): record exploratory-evaluation consensus in CONTEXT.md and ADR-0003
Compare 5 commits »
solahqb pushed to main at solahqb/AgentEvalTool 2026-08-03 07:24:23 +00:00
2d7679a6bb docs(release): freeze v0.8「比」period comparison milestone as 0.8.0
5ecb30876e style(tests): ruff 全量清理 — 49 项修复,backend 与 tests 全绿
14b09e1ac6 feat(campaigns): 周期对比纳入 Markdown 导出,活动导出排版重优化
dd3b9a5e91 refactor(comparison): 评审修复 — 共享 gateway_chat_client、指标元表、对比区块组件化
1e55a21649 feat(comparison): v0.8 周期对比 — 计划指纹自动基线配对、机械指标 diff 与 LLM 演进叙述
Compare 19 commits »