feat: enhance research workflow with user confirmations

- Add AskUserQuestion for Step 1 framework confirmation
- Add TOC summary field selection in report generation
- Add detail_level hierarchy (brief -> moderate -> detailed)
- Add Chinese output requirement for JSON values
- Add long text formatting rules

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
This commit is contained in:
jilinchen
2025-12-31 11:45:24 +08:00
co-authored by Claude Opus 4.5
parent 5eaf859274
commit 11a51a8d79
4 changed files with 276 additions and 251 deletions
+92 -89
View File
@@ -1,140 +1,143 @@
---
allowed-tools: Read, Write, Glob, WebSearch, Task, AskUserQuestion
description: Conduct preliminary research on a topic and generate a research outline. Use for academic research, benchmark research, technology selection, etc.
description: 对目标话题进行初步调研,生成调研outline。用于学术调研、benchmark调研、技术选型等场景。
---
# Research Skill - Preliminary Research
# Research Skill - 初步调研
## Trigger
## 触发方式
`/research <topic>`
## Workflow
## 执行流程
### Step 1: Generate Initial Framework from Model Knowledge
Based on the topic, use model's existing knowledge to generate:
- Main research objects/items list in this domain
- Suggested research field framework
### Step 1: 模型内部知识生成初步框架
基于topic,利用模型已有知识生成:
- 该领域的主要研究对象/items列表
- 建议的调研字段框架
Output {step1_output}.
输出{step1_output},使用AskUserQuestion确认:
- items列表是否需要增减?
- 字段框架是否满足需求?
### Step 2: Web Search Supplement
Use AskUserQuestion to ask for time range (e.g., last 6 months, since 2024, unlimited).
### Step 2: Web Search补充
使用AskUserQuestion询问时间范围(如:最近6个月、2024年至今、不限)。
**Parameter Retrieval**:
- `{topic}`: User input research topic
- `{YYYY-MM-DD}`: Current date
- `{step1_output}`: Complete output from Step 1
- `{time_range}`: User specified time range
**参数获取**
- `{topic}`: 用户输入的调研话题
- `{YYYY-MM-DD}`: 当前日期
- `{step1_output}`: Step 1生成的完整输出内容
- `{time_range}`: 用户指定的时间范围
**Hard Constraint**: The following prompt must be strictly reproduced, only replacing variables in {xxx}, do not modify structure or wording.
**硬约束**:以下prompt必须严格复述,仅替换{xxx}中的变量,禁止改写结构或措辞。
Launch 1 web-search-agent (background), **Prompt Template**:
启动1个web-search-agent(后台),**Prompt模板**
```python
prompt = f"""## Task
Research topic: {topic}
Current date: {YYYY-MM-DD}
prompt = f"""## 任务
调研话题: {topic}
当前日期: {YYYY-MM-DD}
Based on the following initial framework, supplement latest items and recommended research fields.
基于以下初步框架,补充最新items和推荐调研字段。
## Existing Framework
## 已有框架
{step1_output}
## Goals
1. Verify if existing items are missing important objects
2. Supplement items based on missing objects
3. Continue searching for {topic} related items within {time_range} and supplement
4. Supplement new fields
## 目标
1. 验证已有items是否遗漏重要对象
2. 根据遗漏对象进行补充items
3. 继续搜索{topic}相关且{time_range}内的items并补充
4. 补充新fields
## Output Requirements
Return structured results directly (do not write files):
## 输出要求
直接返回结构化结果(不写文件):
### Supplementary Items
- item_name: Brief explanation (why it should be added)
### 补充Items
- item_name: 简要说明(为什么应该加入)
...
### Recommended Supplementary Fields
- field_name: Field description (why this dimension is needed)
### 推荐补充字段
- field_name: 字段描述(为什么需要这个维度)
...
### Sources
- [Source1](url1)
- [Source2](url2)
### 信息来源
- [来源1](url1)
- [来源2](url2)
"""
```
**One-shot Example** (assuming researching AI Coding History):
**One-shot示例**(假设调研AI Coding发展史):
```
## Task
Research topic: AI Coding History
Current date: 2025-12-30
## 任务
调研话题: AI Coding 发展史
当前日期: 2025-12-30
Based on the following initial framework, supplement latest items and recommended research fields.
基于以下初步框架,补充最新items和推荐调研字段。
## Existing Framework
### Items List
1. GitHub Copilot: Developed by Microsoft/GitHub, first mainstream AI coding assistant
2. Cursor: AI-first IDE, based on VSCode
## 已有框架
### Items列表
1. GitHub Copilot: Microsoft/GitHub开发,首个主流AI编程助手
2. Cursor: AI-first IDE,基于VSCode
...
### Field Framework
- Basic Info: name, release_date, company
- Technical Features: underlying_model, context_window
### 字段框架
- 基本信息: name, release_date, company
- 技术特性: underlying_model, context_window
...
## Goals
1. Verify if existing items are missing important objects
2. Supplement items based on missing objects
3. Continue searching for AI Coding History related items within since 2024 and supplement
4. Supplement new fields
## 目标
1. 验证已有items是否遗漏重要对象
2. 根据遗漏对象进行补充items
3. 继续搜索AI Coding 发展史相关且2024年至今内的items并补充
4. 补充新fields
## Output Requirements
Return structured results directly (do not write files):
## 输出要求
直接返回结构化结果(不写文件):
### Supplementary Items
- item_name: Brief explanation (why it should be added)
### 补充Items
- item_name: 简要说明(为什么应该加入)
...
### Recommended Supplementary Fields
- field_name: Field description (why this dimension is needed)
### 推荐补充字段
- field_name: 字段描述(为什么需要这个维度)
...
### Sources
- [Source1](url1)
- [Source2](url2)
### 信息来源
- [来源1](url1)
- [来源2](url2)
```
### Step 3: Ask User for Existing Fields
Use AskUserQuestion to ask if user has existing field definition file, if so read and merge.
### Step 3: 询问用户已有字段
使用AskUserQuestion询问用户是否有已定义的字段文件,如有则读取并合并。
### Step 4: Generate Outline (Separate Files)
Merge {step1_output}, {step2_output} and user's existing fields, generate two files:
### Step 4: 生成Outline(分离文件)
合并{step1_output}{step2_output}和用户已有字段,生成两个文件:
**outline.yaml** (items + config):
- topic: Research topic
- items: Research objects list
**outline.yaml**items + 配置):
- topic: 调研主题
- items: 调研对象列表
- execution:
- batch_size: Number of parallel agents (confirm with AskUserQuestion)
- items_per_agent: Items per agent (confirm with AskUserQuestion)
- output_dir: Results output directory (default: ./results)
- batch_size: 并行agent数量(需AskUserQuestion确认)
- items_per_agent: 每个agent调研项目数(需AskUserQuestion确认)
- output_dir: 结果输出目录(默认./results
**fields.yaml** (field definitions):
- Field categories and definitions
- Each field's name, description, detail_level
- uncertain: Uncertain fields list (reserved field, auto-filled in deep phase)
**fields.yaml**(字段定义):
- 字段分类和定义
- 每个字段的namedescriptiondetail_level
- detail_level分层:极简 → 简要 → 详细
- uncertain: 不确定字段列表(保留字段,deep阶段自动填充)
### Step 5: Output and Confirm
- Create directory: `./{topic_slug}/`
- Save: `outline.yaml` and `fields.yaml`
- Show to user for confirmation
### Step 5: 输出并确认
- 创建目录: `./{topic_slug}/`
- 保存: `outline.yaml` `fields.yaml`
- 展示给用户确认
## Output Path
## 输出路径
```
{current_working_directory}/{topic_slug}/
├── outline.yaml # items list + execution config
└── fields.yaml # field definitions
{当前工作目录}/{topic_slug}/
├── outline.yaml # items列表 + execution配置
└── fields.yaml # 字段定义
```
## Follow-up Commands
- `/research-add-items` - Supplement items
- `/research-add-fields` - Supplement fields
- `/research-deep` - Start deep research
## 后续命令
- `/research-add-items` - 补充items
- `/research-add-fields` - 补充字段
- `/research-deep` - 开始深度调研
+64 -62
View File
@@ -1,96 +1,98 @@
---
description: Read research outline, launch independent agent for each item for deep research. Disable task output.
description: 读取调研outline,为每个item启动独立agent进行深度调研。禁用task output
allowed-tools: Bash, Read, Write, Glob, WebSearch, Task
---
# Research Deep - Deep Research
# Research Deep - 深度调研
## Trigger
## 触发方式
`/research-deep`
## Workflow
## 执行流程
### Step 1: Auto-locate Outline
Find `*/outline.yaml` file in current working directory, read items list, execution config (including items_per_agent).
### Step 1: 自动定位Outline
在当前工作目录查找 `*/outline.yaml` 文件,读取items列表、execution配置(含items_per_agent)。
### Step 2: Resume Check
- Check completed JSON files in output_dir
- Skip completed items
### Step 2: 断点续传检查
- 检查output_dir下已完成的JSON文件
- 跳过已完成的items
### Step 3: Batch Execution
- Batch by batch_size (need user approval before next batch)
- Each agent handles items_per_agent items
- Launch web-search-agent (background parallel, disable task output)
### Step 3: 分批执行
- batch_size分批(完成一批需要得到用户同意才可进行下一批)
- 每个agent负责items_per_agent个项目
- 启动web-search-agent(后台并行,禁用task output
**Parameter Retrieval**:
- `{topic}`: topic field from outline.yaml
- `{item_name}`: item's name field
- `{item_related_info}`: item's complete yaml content (name + category + description etc.)
- `{output_dir}`: execution.output_dir from outline.yaml (default: ./results)
- `{fields_path}`: absolute path to {topic}/fields.yaml
- `{output_path}`: absolute path to {output_dir}/{item_name}.json
**参数获取**
- `{topic}`: outline.yaml中的topic字段
- `{item_name}`: itemname字段
- `{item_related_info}`: item的完整yaml内容(name + category + description等)
- `{output_dir}`: outline.yaml中execution.output_dir(默认./results
- `{fields_path}`: {topic}/fields.yaml的绝对路径
- `{output_path}`: {output_dir}/{item_name}.json的绝对路径
**Hard Constraint**: The following prompt must be strictly reproduced, only replacing variables in {xxx}, do not modify structure or wording.
**硬约束**:以下prompt必须严格复述,仅替换{xxx}中的变量,禁止改写结构或措辞。
**Prompt Template**:
**Prompt模板**
```python
prompt = f"""## Task
Research {item_related_info}, output structured JSON to {output_path}
prompt = f"""## 任务
调研 {item_related_info},输出结构化JSON {output_path}
## Field Definitions
Read {fields_path} to get all field definitions
## 字段定义
读取 {fields_path} 获取所有字段定义
## Output Requirements
1. Output JSON according to fields defined in fields.yaml
2. Mark uncertain field values with [uncertain]
3. Add uncertain array at the end of JSON, listing all uncertain field names
## 输出要求
1. 按fields.yaml定义的字段输出JSON
2. 不确定的字段值标注[不确定]
3. JSON末尾添加uncertain数组,列出所有不确定的字段名
4. 所有字段值必须使用中文输出(调研过程可用英文,但最终JSON值为中文)
## Output Path
## 输出路径
{output_path}
## Validation
After completing JSON output, run validation script to ensure complete field coverage:
## 验证
完成JSON输出后,运行验证脚本确保字段完整覆盖:
python ~/.claude/commands/research/validate_json.py -f {fields_path} -j {output_path}
Task is complete only after validation passes.
验证通过后才算完成任务。
"""
```
**One-shot Example** (assuming researching GitHub Copilot):
**One-shot示例**(假设调研GitHub Copilot):
```
## Task
Research name: GitHub Copilot
category: International Product
description: Developed by Microsoft/GitHub, first mainstream AI coding assistant, ~40% market share, output structured JSON to /home/weizhena/AIcoding/aicoding-history/results/GitHub_Copilot.json
## 任务
调研 name: GitHub Copilot
category: 国际产品
description: Microsoft/GitHub开发,首个主流AI编程助手,市场份额约40%,输出结构化JSON /home/weizhena/AIcoding/aicoding-history/results/GitHub_Copilot.json
## Field Definitions
Read /home/weizhena/AIcoding/aicoding-history/fields.yaml to get all field definitions
## 字段定义
读取 /home/weizhena/AIcoding/aicoding-history/fields.yaml 获取所有字段定义
## Output Requirements
1. Output JSON according to fields defined in fields.yaml
2. Mark uncertain field values with [uncertain]
3. Add uncertain array at the end of JSON, listing all uncertain field names
## 输出要求
1. 按fields.yaml定义的字段输出JSON
2. 不确定的字段值标注[不确定]
3. JSON末尾添加uncertain数组,列出所有不确定的字段名
4. 所有字段值必须使用中文输出(调研过程可用英文,但最终JSON值为中文)
## Output Path
## 输出路径
/home/weizhena/AIcoding/aicoding-history/results/GitHub_Copilot.json
## Validation
After completing JSON output, run validation script to ensure complete field coverage:
## 验证
完成JSON输出后,运行验证脚本确保字段完整覆盖:
python ~/.claude/commands/research/validate_json.py -f /home/weizhena/AIcoding/aicoding-history/fields.yaml -j /home/weizhena/AIcoding/aicoding-history/results/GitHub_Copilot.json
Task is complete only after validation passes.
验证通过后才算完成任务。
```
### Step 4: Wait and Monitor
- Wait for current batch to complete
- Launch next batch
- Display progress
### Step 4: 等待与监控
- 等待当前批次完成
- 启动下一批
- 显示进度
### Step 5: Summary Report
After all complete, output:
- Completion count
- Failed/uncertain marked items
- Output directory
### Step 5: 汇总报告
全部完成后输出:
- 完成数量
- 失败/不确定标记的items
- 输出目录
## Agent Config
- Background execution: Yes
- Task Output: Disabled (agent has explicit output file when complete)
- Resume support: Yes
## Agent配置
- 后台执行: 是
- Task Output: 禁用(agent完成时有明确输出文件)
- 断点续传: 是
+69 -49
View File
@@ -1,71 +1,91 @@
---
description: Summarize deep research results into markdown report, cover all fields, skip uncertain values.
description: 将deep调研结果汇总为markdown报告,覆盖所有字段,跳过不确定值。
allowed-tools: Read, Write, Glob, Bash
---
# Research Report - Summary Report
# Research Report - 汇总报告
## Trigger
## 触发方式
`/research-report`
## Workflow
## 执行流程
### Step 1: Locate Results Directory
Find `*/outline.yaml` in current working directory, read topic and output_dir config.
### Step 1: 定位结果目录
在当前工作目录查找 `*/outline.yaml`,读取topic和output_dir配置。
### Step 2: Generate Python Conversion Script
Generate `generate_report.py` in `{topic}/` directory, script requirements:
- Read all JSON from output_dir
- Read fields.yaml to get field structure
- Cover all field values from each JSON
- Skip fields with values containing [uncertain]
- Skip fields listed in uncertain array
- Generate markdown report format: Table of contents (with anchor links) + Detailed content (by field category)
- Save to `{topic}/report.md`
### Step 2: 扫描可选摘要字段
读取所有JSON结果,提取适合在目录中显示的字段(数值型、简短指标),例如:
- github_stars
- google_scholar_cites
- swe_bench_score
- user_scale
- valuation
- release_date
#### Script Technical Requirements (Must Follow)
使用AskUserQuestion询问用户:
- 目录中除了item名称外,还需要显示哪些字段?
- 提供动态选项列表(基于实际JSON中存在的字段)
**1. JSON Structure Compatibility**
Support two JSON structures:
- Flat structure: Fields directly at top level `{"name": "xxx", "release_date": "xxx"}`
- Nested structure: Fields in category sub-dict `{"basic_info": {"name": "xxx"}, "technical_features": {...}}`
### Step 3: 生成Python转换脚本
`{topic}/` 目录下生成 `generate_report.py`,脚本要求:
- 读取output_dir下所有JSON
- 读取fields.yaml获取字段结构
- 覆盖每个JSON的所有字段值
- 跳过值包含[不确定]的字段
- 跳过uncertain数组中列出的字段
- 生成markdown报告格式:目录(带锚点跳转+用户选择的摘要字段)+ 详细内容(按字段分类)
- 保存到 `{topic}/report.md`
Field lookup order: Top level -> category mapping key -> Traverse all nested dicts
**目录格式要求**
- 必须包含每一个item
- 每个item显示:序号、名称(锚点链接)、用户选择的摘要字段
- 示例:`1. [GitHub Copilot](#github-copilot) - Stars: 10k | Score: 85%`
**2. Category Multi-language Mapping**
fields.yaml category names and JSON keys can be any combination (CN-CN, CN-EN, EN-CN, EN-EN). Must establish bidirectional mapping:
#### 脚本技术要点(必须遵循)
**1. JSON结构兼容**
支持两种JSON结构:
- 扁平结构:字段直接在顶层 `{"name": "xxx", "release_date": "xxx"}`
- 嵌套结构:字段在category子dict中 `{"basic_info": {"name": "xxx"}, "technical_features": {...}}`
字段查找顺序:顶层 -> category映射key -> 遍历所有嵌套dict
**2. Category多语言映射**
fields.yaml的category名与JSON的key可能是任意组合(中中、中英、英中、英英)。必须建立双向映射:
```python
CATEGORY_MAPPING = {
"Basic Info": ["basic_info", "Basic Info"],
"Technical Features": ["technical_features", "technical_characteristics", "Technical Features"],
"Performance Metrics": ["performance_metrics", "performance", "Performance Metrics"],
"Milestone Significance": ["milestone_significance", "milestones", "Milestone Significance"],
"Business Info": ["business_info", "commercial_info", "Business Info"],
"Competition & Ecosystem": ["competition_ecosystem", "competition", "Competition & Ecosystem"],
"History": ["history", "History"],
"Market Positioning": ["market_positioning", "market", "Market Positioning"],
"基本信息": ["basic_info", "基本信息"],
"技术特性": ["technical_features", "technical_characteristics", "技术特性"],
"性能指标": ["performance_metrics", "performance", "性能指标"],
"里程碑意义": ["milestone_significance", "milestones", "里程碑意义"],
"商业信息": ["business_info", "commercial_info", "商业信息"],
"竞争与生态": ["competition_ecosystem", "competition", "竞争与生态"],
"历史沿革": ["history", "历史沿革"],
"市场定位": ["market_positioning", "market", "市场定位"],
}
```
**3. Complex Value Formatting**
- list of dicts (e.g., key_events, funding_history): Format each dict as one line, separate kv with ` | `
- Normal list: Short lists joined with comma, long lists displayed with line breaks
- Nested dict: Recursive formatting, display with semicolon or line breaks
**3. 复杂值格式化**
- list of dicts(如key_events, funding_history):每个dict格式化为一行,用` | `分隔kv
- 普通list:短列表用逗号连接,长列表换行显示
- 嵌套dict:递归格式化,用分号或换行显示
- 长文本字符串(超过100字符):添加换行符`<br>`或使用blockquote格式,提高可读性
**4. Extra Fields Collection**
Collect fields that exist in JSON but not defined in fields.yaml, put in "Other Info" category. Note to filter:
- Internal fields: `_source_file`, `uncertain`
- Nested structure top-level keys: `basic_info`, `technical_features` etc.
**4. 额外字段收集**
收集JSON中有但fields.yaml中没定义的字段,放入"其他信息"分类。注意过滤:
- 内部字段:`_source_file`, `uncertain`
- 嵌套结构顶级key`basic_info`, `technical_features`
- `uncertain_fields`列表:需要逐行显示每个字段名,不要压缩成一行
**5. Uncertain Value Skipping**
Skip conditions:
- Field value contains `[uncertain]` string
- Field name is in `uncertain` array
- Field value is None or empty string
**5. 不确定值跳过**
跳过条件:
- 字段值包含`[不确定]`字符串
- 字段名在`uncertain`数组中
- 字段值为None或空字符串
### Step 3: Execute Script
Run `python {topic}/generate_report.py`
### Step 4: 执行脚本
运行 `python {topic}/generate_report.py`
## Output
- `{topic}/generate_report.py` - Conversion script
- `{topic}/report.md` - Summary report
## 输出
- `{topic}/generate_report.py` - 转换脚本
- `{topic}/report.md` - 汇总报告
+51 -51
View File
@@ -1,8 +1,8 @@
#!/usr/bin/env python3
# -*- coding: utf-8 -*-
"""
JSON Field Validation Script
Validate if JSON files completely cover all fields defined in fields.yaml
JSON字段验证脚本
验证JSON文件是否完整覆盖fields.yaml中定义的所有字段
"""
import json
@@ -11,25 +11,25 @@ import sys
from pathlib import Path
from typing import Dict, List, Set, Any, Tuple
# Category mapping (supports both Chinese and English)
# Category中英文映射
CATEGORY_MAPPING = {
"Basic Info": ["basic_info", "Basic Info"],
"Technical Features": ["technical_features", "technical_characteristics", "Technical Features"],
"Performance Metrics": ["performance_metrics", "performance", "Performance Metrics"],
"Milestone Significance": ["milestone_significance", "milestones", "Milestone Significance"],
"Business Info": ["business_info", "commercial_info", "Business Info"],
"Competition & Ecosystem": ["competition_ecosystem", "competition", "Competition & Ecosystem"],
"History": ["history", "History"],
"Market Positioning": ["market_positioning", "market", "Market Positioning"],
"基本信息": ["basic_info", "基本信息"],
"技术特性": ["technical_features", "technical_characteristics", "技术特性"],
"性能指标": ["performance_metrics", "performance", "性能指标"],
"里程碑意义": ["milestone_significance", "milestones", "里程碑意义"],
"商业信息": ["business_info", "commercial_info", "商业信息"],
"竞争与生态": ["competition_ecosystem", "competition", "竞争与生态"],
"历史沿革": ["history", "历史沿革"],
"市场定位": ["market_positioning", "market", "市场定位"],
}
def load_fields_yaml(fields_path: Path) -> Tuple[Set[str], Set[str], Dict[str, str]]:
"""
Load fields.yaml, returns:
- all_fields: Set of all field names
- required_fields: Set of field names where required=true
- field_categories: Mapping from field name to category
加载fields.yaml,返回:
- all_fields: 所有字段名集合
- required_fields: required=true的字段名集合
- field_categories: 字段名到类别的映射
"""
with open(fields_path, 'r', encoding='utf-8') as f:
data = yaml.safe_load(f)
@@ -52,12 +52,12 @@ def load_fields_yaml(fields_path: Path) -> Tuple[Set[str], Set[str], Dict[str, s
def extract_json_fields(data: Dict, category_mapping: Dict = None) -> Set[str]:
"""
Extract all field names from JSON (supports flat and nested structures)
从JSON中提取所有字段名(支持扁平和嵌套结构)
"""
if category_mapping is None:
category_mapping = CATEGORY_MAPPING
# Get all possible nested keys
# 获取所有可能的嵌套key
nested_keys = set()
for keys in category_mapping.values():
nested_keys.update(keys)
@@ -66,10 +66,10 @@ def extract_json_fields(data: Dict, category_mapping: Dict = None) -> Set[str]:
def collect_fields(d: Dict, is_top_level: bool = True):
for k, v in d.items():
# Skip internal fields
# 跳过内部字段
if k in {"_source_file", "uncertain"}:
continue
# If it's a nested structure top-level key, recurse into it
# 如果是嵌套结构的顶级key,递归进入
if is_top_level and k in nested_keys:
if isinstance(v, dict):
collect_fields(v, is_top_level=False)
@@ -85,27 +85,27 @@ def extract_json_fields(data: Dict, category_mapping: Dict = None) -> Set[str]:
def validate_json(json_path: Path, all_fields: Set[str], required_fields: Set[str],
field_categories: Dict[str, str]) -> Dict:
"""
Validate a single JSON file
Returns validation result dictionary
验证单个JSON文件
返回验证结果字典
"""
with open(json_path, 'r', encoding='utf-8') as f:
data = json.load(f)
json_fields = extract_json_fields(data)
# Calculate coverage
# 计算覆盖情况
covered = all_fields & json_fields
missing = all_fields - json_fields
extra = json_fields - all_fields
# Categorize missing fields
# 分类缺失字段
missing_required = missing & required_fields
missing_optional = missing - required_fields
# Group missing fields by category
# 按类别分组缺失字段
missing_by_category = {}
for field in missing:
cat = field_categories.get(field, "Unknown")
cat = field_categories.get(field, "未知")
if cat not in missing_by_category:
missing_by_category[cat] = []
missing_by_category[cat].append(field)
@@ -121,65 +121,65 @@ def validate_json(json_path: Path, all_fields: Set[str], required_fields: Set[st
"missing_optional": list(missing_optional),
"missing_by_category": missing_by_category,
"extra_fields": list(extra),
"valid": len(missing_required) == 0, # Valid if all required fields are covered
"valid": len(missing_required) == 0, # required字段全覆盖则valid
}
def print_result(result: Dict, verbose: bool = True):
"""Print validation result"""
"""打印验证结果"""
status = "PASS" if result["valid"] else "FAIL"
print(f"\n{'='*60}")
print(f"[{status}] {result['file']}")
print(f"{'='*60}")
print(f"Coverage: {result['coverage_rate']:.1f}% ({result['covered']}/{result['total_defined']})")
print(f"覆盖率: {result['coverage_rate']:.1f}% ({result['covered']}/{result['total_defined']})")
if result["missing_required"]:
print(f"\n[ERROR] Missing required fields ({len(result['missing_required'])}):")
print(f"\n[ERROR] 缺失必需字段 ({len(result['missing_required'])}):")
for field in result["missing_required"]:
print(f" - {field}")
if verbose and result["missing_optional"]:
print(f"\n[WARN] Missing optional fields ({len(result['missing_optional'])}):")
print(f"\n[WARN] 缺失可选字段 ({len(result['missing_optional'])}):")
for cat, fields in result["missing_by_category"].items():
optional_fields = [f for f in fields if f not in result["missing_required"]]
if optional_fields:
print(f" [{cat}]: {', '.join(optional_fields)}")
if verbose and result["extra_fields"]:
print(f"\n[INFO] Extra fields ({len(result['extra_fields'])}):")
print(f"\n[INFO] 额外字段 ({len(result['extra_fields'])}):")
print(f" {', '.join(result['extra_fields'][:10])}")
if len(result["extra_fields"]) > 10:
print(f" ... and {len(result['extra_fields']) - 10} more fields")
print(f" ... {len(result['extra_fields']) - 10} 个其他字段")
def main():
"""Main function"""
"""主函数"""
import argparse
parser = argparse.ArgumentParser(description="Validate if JSON files cover all fields defined in fields.yaml")
parser.add_argument("--fields", "-f", type=str, help="Path to fields.yaml", default="fields.yaml")
parser.add_argument("--json", "-j", type=str, nargs="*", help="JSON file paths to validate")
parser.add_argument("--dir", "-d", type=str, help="JSON files directory", default="results")
parser.add_argument("--quiet", "-q", action="store_true", help="Show summary only")
parser = argparse.ArgumentParser(description="验证JSON文件是否覆盖fields.yaml定义的所有字段")
parser.add_argument("--fields", "-f", type=str, help="fields.yaml路径", default="fields.yaml")
parser.add_argument("--json", "-j", type=str, nargs="*", help="要验证的JSON文件路径")
parser.add_argument("--dir", "-d", type=str, help="JSON文件目录", default="results")
parser.add_argument("--quiet", "-q", action="store_true", help="只显示摘要")
args = parser.parse_args()
# Locate fields.yaml
# 定位fields.yaml
fields_path = Path(args.fields)
if not fields_path.exists():
# Try to find in current and parent directory
# 尝试在当前目录和父目录查找
for p in [Path.cwd() / "fields.yaml", Path.cwd().parent / "fields.yaml"]:
if p.exists():
fields_path = p
break
if not fields_path.exists():
print(f"[ERROR] fields.yaml not found: {fields_path}")
print(f"[ERROR] fields.yaml不存在: {fields_path}")
sys.exit(1)
print(f"Field definitions file: {fields_path}")
print(f"字段定义文件: {fields_path}")
all_fields, required_fields, field_categories = load_fields_yaml(fields_path)
print(f"Total fields: {len(all_fields)} (required: {len(required_fields)}, optional: {len(all_fields) - len(required_fields)})")
print(f"总字段数: {len(all_fields)} (必需: {len(required_fields)}, 可选: {len(all_fields) - len(required_fields)})")
# Collect JSON files
# 收集JSON文件
json_files = []
if args.json:
json_files = [Path(p) for p in args.json]
@@ -189,27 +189,27 @@ def main():
json_files = sorted(json_dir.glob("*.json"))
if not json_files:
print(f"[WARN] No JSON files found")
print(f"[WARN] 未找到JSON文件")
sys.exit(0)
# Validate each file
# 验证每个文件
results = []
for json_path in json_files:
if not json_path.exists():
print(f"[WARN] File not found: {json_path}")
print(f"[WARN] 文件不存在: {json_path}")
continue
result = validate_json(json_path, all_fields, required_fields, field_categories)
results.append(result)
print_result(result, verbose=not args.quiet)
# Summary
# 汇总
print(f"\n{'='*60}")
print("Summary")
print("汇总")
print(f"{'='*60}")
passed = sum(1 for r in results if r["valid"])
avg_coverage = sum(r["coverage_rate"] for r in results) / len(results) if results else 0
print(f"Validation passed: {passed}/{len(results)}")
print(f"Average coverage: {avg_coverage:.1f}%")
print(f"验证通过: {passed}/{len(results)}")
print(f"平均覆盖率: {avg_coverage:.1f}%")
if passed < len(results):
sys.exit(1)