• Joined on 2026-06-26
solahqb commented on issue solahqb/AgentEvalTool#14 2026-08-25 02:53:51 +00:00
定义评估维度清单

Resolution

6 个评估维度及量化方式

solahqb closed issue solahqb/AgentEvalTool#14 2026-08-25 02:53:51 +00:00
定义评估维度清单
solahqb closed issue solahqb/AgentEvalTool#16 2026-08-24 17:51:16 +00:00
当前平台能力盘点
solahqb commented on issue solahqb/AgentEvalTool#16 2026-08-24 17:51:16 +00:00
当前平台能力盘点

Resolution

能力盘点完成,详见 [research/platform-capability-inventory.md](https://git.solahqb22.cn/solahqb/AgentEvalTool/blob/research/platform-capability/research/platform-capability-

solahqb opened issue solahqb/AgentEvalTool#15 2026-08-24 17:47:37 +00:00
定义报告消费者与决策场景
solahqb opened issue solahqb/AgentEvalTool#16 2026-08-24 17:47:37 +00:00
当前平台能力盘点
solahqb opened issue solahqb/AgentEvalTool#14 2026-08-24 17:47:24 +00:00
定义评估维度清单
solahqb opened issue solahqb/AgentEvalTool#13 2026-08-24 17:46:48 +00:00
Wayfinder: AgentEvalTool 价值评估与决策
solahqb closed pull request solahqb/AgentEvalTool#12 2026-08-24 17:34:03 +00:00
docs(readme): 添加 v1.3.0 性能基准实测数据
solahqb created pull request solahqb/AgentEvalTool#12 2026-08-24 17:26:44 +00:00
docs(readme): 添加 v1.3.0 性能基准实测数据
solahqb created branch docs/readme-benchmark-v1.3 in solahqb/AgentEvalTool 2026-08-24 17:14:59 +00:00
solahqb pushed to docs/readme-benchmark-v1.3 at solahqb/AgentEvalTool 2026-08-24 17:14:59 +00:00
b1a09f02b7 docs(readme): 添加 v1.3.0 性能基准实测数据
solahqb pushed to main at solahqb/AgentEvalTool 2026-08-24 15:39:28 +00:00
0d074e5465 test(runs): 取消测试扩至 30 case 适配并发执行
bd6efc3709 chore: bump version to 1.3.0
9f5da70c36 docs: 归档 2026-08-24 全面代码库健康审查报告
bf1ec16ef6 perf(eval): case/rule/campaign 并发执行
ca208232c7 perf(db): WAL 模式 + 性能索引 + N+1 查询消除
Compare 8 commits »
solahqb pushed to main at solahqb/AgentEvalTool 2026-08-23 21:58:37 +00:00
9588dcbfda docs: 归档 2026-08-24 架构审查报告(Phase 5)
4534df7e7a test: 五个零覆盖组件补全测试(Phase 4.23)
eca7ccebe2 refactor(api): api.ts 拆分为 10 个域模块(Phase 4.22)
f458897ab5 refactor(read): 前端泛化 ReadSlot 资源接缝(Phase 4.18-21)
7eae6de52d refactor(evaluation/storage): 结算统一与 repository 拆分(Phase 2 + 3)
Compare 6 commits »
solahqb pushed to main at solahqb/AgentEvalTool 2026-08-23 18:01:43 +00:00
913dc9ae86 perf: address remaining heuristic issues from code review
solahqb pushed to main at solahqb/AgentEvalTool 2026-08-23 17:54:06 +00:00
5a81c570c0 style: fix ruff whitespace warnings
09ff2ed123 refactor(intelligent-eval): reduce nesting complexity in supplement_decision_logs
da7dd434dd perf(intelligent-eval): 修复 N+1 查询和参数名混淆
Compare 3 commits »
solahqb pushed to main at solahqb/AgentEvalTool 2026-08-23 17:48:22 +00:00
84627a3c6a refactor(intelligent-eval): 消除 lifecycle.py 和 scheduler.py 中的重复延迟导入
solahqb pushed to main at solahqb/AgentEvalTool 2026-08-23 17:38:39 +00:00
876d75f9ed refactor(intelligent-eval): 将 _get_raw 改为公开方法 get_including_deleted
solahqb pushed to main at solahqb/AgentEvalTool 2026-08-23 05:33:24 +00:00
7db75be707 fix(ui): 评估列表默认分页改为 20,符合 ADR-0005
solahqb pushed tag v1.2.0 to solahqb/AgentEvalTool 2026-08-21 09:54:05 +00:00