feat: enhance research workflow with user confirmations
- Add AskUserQuestion for Step 1 framework confirmation - Add TOC summary field selection in report generation - Add detail_level hierarchy (brief -> moderate -> detailed) - Add Chinese output requirement for JSON values - Add long text formatting rules 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 4.5
parent
5eaf859274
commit
11a51a8d79
+92
-89
@@ -1,140 +1,143 @@
|
|||||||
---
|
---
|
||||||
allowed-tools: Read, Write, Glob, WebSearch, Task, AskUserQuestion
|
allowed-tools: Read, Write, Glob, WebSearch, Task, AskUserQuestion
|
||||||
description: Conduct preliminary research on a topic and generate a research outline. Use for academic research, benchmark research, technology selection, etc.
|
description: 对目标话题进行初步调研,生成调研outline。用于学术调研、benchmark调研、技术选型等场景。
|
||||||
---
|
---
|
||||||
|
|
||||||
# Research Skill - Preliminary Research
|
# Research Skill - 初步调研
|
||||||
|
|
||||||
## Trigger
|
## 触发方式
|
||||||
`/research <topic>`
|
`/research <topic>`
|
||||||
|
|
||||||
## Workflow
|
## 执行流程
|
||||||
|
|
||||||
### Step 1: Generate Initial Framework from Model Knowledge
|
### Step 1: 模型内部知识生成初步框架
|
||||||
Based on the topic, use model's existing knowledge to generate:
|
基于topic,利用模型已有知识生成:
|
||||||
- Main research objects/items list in this domain
|
- 该领域的主要研究对象/items列表
|
||||||
- Suggested research field framework
|
- 建议的调研字段框架
|
||||||
|
|
||||||
Output {step1_output}.
|
输出{step1_output},使用AskUserQuestion确认:
|
||||||
|
- items列表是否需要增减?
|
||||||
|
- 字段框架是否满足需求?
|
||||||
|
|
||||||
### Step 2: Web Search Supplement
|
### Step 2: Web Search补充
|
||||||
Use AskUserQuestion to ask for time range (e.g., last 6 months, since 2024, unlimited).
|
使用AskUserQuestion询问时间范围(如:最近6个月、2024年至今、不限)。
|
||||||
|
|
||||||
**Parameter Retrieval**:
|
**参数获取**:
|
||||||
- `{topic}`: User input research topic
|
- `{topic}`: 用户输入的调研话题
|
||||||
- `{YYYY-MM-DD}`: Current date
|
- `{YYYY-MM-DD}`: 当前日期
|
||||||
- `{step1_output}`: Complete output from Step 1
|
- `{step1_output}`: Step 1生成的完整输出内容
|
||||||
- `{time_range}`: User specified time range
|
- `{time_range}`: 用户指定的时间范围
|
||||||
|
|
||||||
**Hard Constraint**: The following prompt must be strictly reproduced, only replacing variables in {xxx}, do not modify structure or wording.
|
**硬约束**:以下prompt必须严格复述,仅替换{xxx}中的变量,禁止改写结构或措辞。
|
||||||
|
|
||||||
Launch 1 web-search-agent (background), **Prompt Template**:
|
启动1个web-search-agent(后台),**Prompt模板**:
|
||||||
```python
|
```python
|
||||||
prompt = f"""## Task
|
prompt = f"""## 任务
|
||||||
Research topic: {topic}
|
调研话题: {topic}
|
||||||
Current date: {YYYY-MM-DD}
|
当前日期: {YYYY-MM-DD}
|
||||||
|
|
||||||
Based on the following initial framework, supplement latest items and recommended research fields.
|
基于以下初步框架,补充最新items和推荐调研字段。
|
||||||
|
|
||||||
## Existing Framework
|
## 已有框架
|
||||||
{step1_output}
|
{step1_output}
|
||||||
|
|
||||||
## Goals
|
## 目标
|
||||||
1. Verify if existing items are missing important objects
|
1. 验证已有items是否遗漏重要对象
|
||||||
2. Supplement items based on missing objects
|
2. 根据遗漏对象进行补充items
|
||||||
3. Continue searching for {topic} related items within {time_range} and supplement
|
3. 继续搜索{topic}相关且{time_range}内的items并补充
|
||||||
4. Supplement new fields
|
4. 补充新fields
|
||||||
|
|
||||||
## Output Requirements
|
## 输出要求
|
||||||
Return structured results directly (do not write files):
|
直接返回结构化结果(不写文件):
|
||||||
|
|
||||||
### Supplementary Items
|
### 补充Items
|
||||||
- item_name: Brief explanation (why it should be added)
|
- item_name: 简要说明(为什么应该加入)
|
||||||
...
|
...
|
||||||
|
|
||||||
### Recommended Supplementary Fields
|
### 推荐补充字段
|
||||||
- field_name: Field description (why this dimension is needed)
|
- field_name: 字段描述(为什么需要这个维度)
|
||||||
...
|
...
|
||||||
|
|
||||||
### Sources
|
### 信息来源
|
||||||
- [Source1](url1)
|
- [来源1](url1)
|
||||||
- [Source2](url2)
|
- [来源2](url2)
|
||||||
"""
|
"""
|
||||||
```
|
```
|
||||||
|
|
||||||
**One-shot Example** (assuming researching AI Coding History):
|
**One-shot示例**(假设调研AI Coding发展史):
|
||||||
```
|
```
|
||||||
## Task
|
## 任务
|
||||||
Research topic: AI Coding History
|
调研话题: AI Coding 发展史
|
||||||
Current date: 2025-12-30
|
当前日期: 2025-12-30
|
||||||
|
|
||||||
Based on the following initial framework, supplement latest items and recommended research fields.
|
基于以下初步框架,补充最新items和推荐调研字段。
|
||||||
|
|
||||||
## Existing Framework
|
## 已有框架
|
||||||
### Items List
|
### Items列表
|
||||||
1. GitHub Copilot: Developed by Microsoft/GitHub, first mainstream AI coding assistant
|
1. GitHub Copilot: Microsoft/GitHub开发,首个主流AI编程助手
|
||||||
2. Cursor: AI-first IDE, based on VSCode
|
2. Cursor: AI-first IDE,基于VSCode
|
||||||
...
|
...
|
||||||
|
|
||||||
### Field Framework
|
### 字段框架
|
||||||
- Basic Info: name, release_date, company
|
- 基本信息: name, release_date, company
|
||||||
- Technical Features: underlying_model, context_window
|
- 技术特性: underlying_model, context_window
|
||||||
...
|
...
|
||||||
|
|
||||||
## Goals
|
## 目标
|
||||||
1. Verify if existing items are missing important objects
|
1. 验证已有items是否遗漏重要对象
|
||||||
2. Supplement items based on missing objects
|
2. 根据遗漏对象进行补充items
|
||||||
3. Continue searching for AI Coding History related items within since 2024 and supplement
|
3. 继续搜索AI Coding 发展史相关且2024年至今内的items并补充
|
||||||
4. Supplement new fields
|
4. 补充新fields
|
||||||
|
|
||||||
## Output Requirements
|
## 输出要求
|
||||||
Return structured results directly (do not write files):
|
直接返回结构化结果(不写文件):
|
||||||
|
|
||||||
### Supplementary Items
|
### 补充Items
|
||||||
- item_name: Brief explanation (why it should be added)
|
- item_name: 简要说明(为什么应该加入)
|
||||||
...
|
...
|
||||||
|
|
||||||
### Recommended Supplementary Fields
|
### 推荐补充字段
|
||||||
- field_name: Field description (why this dimension is needed)
|
- field_name: 字段描述(为什么需要这个维度)
|
||||||
...
|
...
|
||||||
|
|
||||||
### Sources
|
### 信息来源
|
||||||
- [Source1](url1)
|
- [来源1](url1)
|
||||||
- [Source2](url2)
|
- [来源2](url2)
|
||||||
```
|
```
|
||||||
|
|
||||||
### Step 3: Ask User for Existing Fields
|
### Step 3: 询问用户已有字段
|
||||||
Use AskUserQuestion to ask if user has existing field definition file, if so read and merge.
|
使用AskUserQuestion询问用户是否有已定义的字段文件,如有则读取并合并。
|
||||||
|
|
||||||
### Step 4: Generate Outline (Separate Files)
|
### Step 4: 生成Outline(分离文件)
|
||||||
Merge {step1_output}, {step2_output} and user's existing fields, generate two files:
|
合并{step1_output}、{step2_output}和用户已有字段,生成两个文件:
|
||||||
|
|
||||||
**outline.yaml** (items + config):
|
**outline.yaml**(items + 配置):
|
||||||
- topic: Research topic
|
- topic: 调研主题
|
||||||
- items: Research objects list
|
- items: 调研对象列表
|
||||||
- execution:
|
- execution:
|
||||||
- batch_size: Number of parallel agents (confirm with AskUserQuestion)
|
- batch_size: 并行agent数量(需AskUserQuestion确认)
|
||||||
- items_per_agent: Items per agent (confirm with AskUserQuestion)
|
- items_per_agent: 每个agent调研项目数(需AskUserQuestion确认)
|
||||||
- output_dir: Results output directory (default: ./results)
|
- output_dir: 结果输出目录(默认./results)
|
||||||
|
|
||||||
**fields.yaml** (field definitions):
|
**fields.yaml**(字段定义):
|
||||||
- Field categories and definitions
|
- 字段分类和定义
|
||||||
- Each field's name, description, detail_level
|
- 每个字段的name、description、detail_level
|
||||||
- uncertain: Uncertain fields list (reserved field, auto-filled in deep phase)
|
- detail_level分层:极简 → 简要 → 详细
|
||||||
|
- uncertain: 不确定字段列表(保留字段,deep阶段自动填充)
|
||||||
|
|
||||||
### Step 5: Output and Confirm
|
### Step 5: 输出并确认
|
||||||
- Create directory: `./{topic_slug}/`
|
- 创建目录: `./{topic_slug}/`
|
||||||
- Save: `outline.yaml` and `fields.yaml`
|
- 保存: `outline.yaml` 和 `fields.yaml`
|
||||||
- Show to user for confirmation
|
- 展示给用户确认
|
||||||
|
|
||||||
## Output Path
|
## 输出路径
|
||||||
```
|
```
|
||||||
{current_working_directory}/{topic_slug}/
|
{当前工作目录}/{topic_slug}/
|
||||||
├── outline.yaml # items list + execution config
|
├── outline.yaml # items列表 + execution配置
|
||||||
└── fields.yaml # field definitions
|
└── fields.yaml # 字段定义
|
||||||
```
|
```
|
||||||
|
|
||||||
## Follow-up Commands
|
## 后续命令
|
||||||
- `/research-add-items` - Supplement items
|
- `/research-add-items` - 补充items
|
||||||
- `/research-add-fields` - Supplement fields
|
- `/research-add-fields` - 补充字段
|
||||||
- `/research-deep` - Start deep research
|
- `/research-deep` - 开始深度调研
|
||||||
|
|||||||
@@ -1,96 +1,98 @@
|
|||||||
---
|
---
|
||||||
description: Read research outline, launch independent agent for each item for deep research. Disable task output.
|
description: 读取调研outline,为每个item启动独立agent进行深度调研。禁用task output。
|
||||||
allowed-tools: Bash, Read, Write, Glob, WebSearch, Task
|
allowed-tools: Bash, Read, Write, Glob, WebSearch, Task
|
||||||
---
|
---
|
||||||
|
|
||||||
# Research Deep - Deep Research
|
# Research Deep - 深度调研
|
||||||
|
|
||||||
## Trigger
|
## 触发方式
|
||||||
`/research-deep`
|
`/research-deep`
|
||||||
|
|
||||||
## Workflow
|
## 执行流程
|
||||||
|
|
||||||
### Step 1: Auto-locate Outline
|
### Step 1: 自动定位Outline
|
||||||
Find `*/outline.yaml` file in current working directory, read items list, execution config (including items_per_agent).
|
在当前工作目录查找 `*/outline.yaml` 文件,读取items列表、execution配置(含items_per_agent)。
|
||||||
|
|
||||||
### Step 2: Resume Check
|
### Step 2: 断点续传检查
|
||||||
- Check completed JSON files in output_dir
|
- 检查output_dir下已完成的JSON文件
|
||||||
- Skip completed items
|
- 跳过已完成的items
|
||||||
|
|
||||||
### Step 3: Batch Execution
|
### Step 3: 分批执行
|
||||||
- Batch by batch_size (need user approval before next batch)
|
- 按batch_size分批(完成一批需要得到用户同意才可进行下一批)
|
||||||
- Each agent handles items_per_agent items
|
- 每个agent负责items_per_agent个项目
|
||||||
- Launch web-search-agent (background parallel, disable task output)
|
- 启动web-search-agent(后台并行,禁用task output)
|
||||||
|
|
||||||
**Parameter Retrieval**:
|
**参数获取**:
|
||||||
- `{topic}`: topic field from outline.yaml
|
- `{topic}`: outline.yaml中的topic字段
|
||||||
- `{item_name}`: item's name field
|
- `{item_name}`: item的name字段
|
||||||
- `{item_related_info}`: item's complete yaml content (name + category + description etc.)
|
- `{item_related_info}`: item的完整yaml内容(name + category + description等)
|
||||||
- `{output_dir}`: execution.output_dir from outline.yaml (default: ./results)
|
- `{output_dir}`: outline.yaml中execution.output_dir(默认./results)
|
||||||
- `{fields_path}`: absolute path to {topic}/fields.yaml
|
- `{fields_path}`: {topic}/fields.yaml的绝对路径
|
||||||
- `{output_path}`: absolute path to {output_dir}/{item_name}.json
|
- `{output_path}`: {output_dir}/{item_name}.json的绝对路径
|
||||||
|
|
||||||
**Hard Constraint**: The following prompt must be strictly reproduced, only replacing variables in {xxx}, do not modify structure or wording.
|
**硬约束**:以下prompt必须严格复述,仅替换{xxx}中的变量,禁止改写结构或措辞。
|
||||||
|
|
||||||
**Prompt Template**:
|
**Prompt模板**:
|
||||||
```python
|
```python
|
||||||
prompt = f"""## Task
|
prompt = f"""## 任务
|
||||||
Research {item_related_info}, output structured JSON to {output_path}
|
调研 {item_related_info},输出结构化JSON到 {output_path}
|
||||||
|
|
||||||
## Field Definitions
|
## 字段定义
|
||||||
Read {fields_path} to get all field definitions
|
读取 {fields_path} 获取所有字段定义
|
||||||
|
|
||||||
## Output Requirements
|
## 输出要求
|
||||||
1. Output JSON according to fields defined in fields.yaml
|
1. 按fields.yaml定义的字段输出JSON
|
||||||
2. Mark uncertain field values with [uncertain]
|
2. 不确定的字段值标注[不确定]
|
||||||
3. Add uncertain array at the end of JSON, listing all uncertain field names
|
3. JSON末尾添加uncertain数组,列出所有不确定的字段名
|
||||||
|
4. 所有字段值必须使用中文输出(调研过程可用英文,但最终JSON值为中文)
|
||||||
|
|
||||||
## Output Path
|
## 输出路径
|
||||||
{output_path}
|
{output_path}
|
||||||
|
|
||||||
## Validation
|
## 验证
|
||||||
After completing JSON output, run validation script to ensure complete field coverage:
|
完成JSON输出后,运行验证脚本确保字段完整覆盖:
|
||||||
python ~/.claude/commands/research/validate_json.py -f {fields_path} -j {output_path}
|
python ~/.claude/commands/research/validate_json.py -f {fields_path} -j {output_path}
|
||||||
Task is complete only after validation passes.
|
验证通过后才算完成任务。
|
||||||
"""
|
"""
|
||||||
```
|
```
|
||||||
|
|
||||||
**One-shot Example** (assuming researching GitHub Copilot):
|
**One-shot示例**(假设调研GitHub Copilot):
|
||||||
```
|
```
|
||||||
## Task
|
## 任务
|
||||||
Research name: GitHub Copilot
|
调研 name: GitHub Copilot
|
||||||
category: International Product
|
category: 国际产品
|
||||||
description: Developed by Microsoft/GitHub, first mainstream AI coding assistant, ~40% market share, output structured JSON to /home/weizhena/AIcoding/aicoding-history/results/GitHub_Copilot.json
|
description: Microsoft/GitHub开发,首个主流AI编程助手,市场份额约40%,输出结构化JSON到 /home/weizhena/AIcoding/aicoding-history/results/GitHub_Copilot.json
|
||||||
|
|
||||||
## Field Definitions
|
## 字段定义
|
||||||
Read /home/weizhena/AIcoding/aicoding-history/fields.yaml to get all field definitions
|
读取 /home/weizhena/AIcoding/aicoding-history/fields.yaml 获取所有字段定义
|
||||||
|
|
||||||
## Output Requirements
|
## 输出要求
|
||||||
1. Output JSON according to fields defined in fields.yaml
|
1. 按fields.yaml定义的字段输出JSON
|
||||||
2. Mark uncertain field values with [uncertain]
|
2. 不确定的字段值标注[不确定]
|
||||||
3. Add uncertain array at the end of JSON, listing all uncertain field names
|
3. JSON末尾添加uncertain数组,列出所有不确定的字段名
|
||||||
|
4. 所有字段值必须使用中文输出(调研过程可用英文,但最终JSON值为中文)
|
||||||
|
|
||||||
## Output Path
|
## 输出路径
|
||||||
/home/weizhena/AIcoding/aicoding-history/results/GitHub_Copilot.json
|
/home/weizhena/AIcoding/aicoding-history/results/GitHub_Copilot.json
|
||||||
|
|
||||||
## Validation
|
## 验证
|
||||||
After completing JSON output, run validation script to ensure complete field coverage:
|
完成JSON输出后,运行验证脚本确保字段完整覆盖:
|
||||||
python ~/.claude/commands/research/validate_json.py -f /home/weizhena/AIcoding/aicoding-history/fields.yaml -j /home/weizhena/AIcoding/aicoding-history/results/GitHub_Copilot.json
|
python ~/.claude/commands/research/validate_json.py -f /home/weizhena/AIcoding/aicoding-history/fields.yaml -j /home/weizhena/AIcoding/aicoding-history/results/GitHub_Copilot.json
|
||||||
Task is complete only after validation passes.
|
验证通过后才算完成任务。
|
||||||
```
|
```
|
||||||
|
|
||||||
### Step 4: Wait and Monitor
|
### Step 4: 等待与监控
|
||||||
- Wait for current batch to complete
|
- 等待当前批次完成
|
||||||
- Launch next batch
|
- 启动下一批
|
||||||
- Display progress
|
- 显示进度
|
||||||
|
|
||||||
### Step 5: Summary Report
|
### Step 5: 汇总报告
|
||||||
After all complete, output:
|
全部完成后输出:
|
||||||
- Completion count
|
- 完成数量
|
||||||
- Failed/uncertain marked items
|
- 失败/不确定标记的items
|
||||||
- Output directory
|
- 输出目录
|
||||||
|
|
||||||
## Agent Config
|
## Agent配置
|
||||||
- Background execution: Yes
|
- 后台执行: 是
|
||||||
- Task Output: Disabled (agent has explicit output file when complete)
|
- Task Output: 禁用(agent完成时有明确输出文件)
|
||||||
- Resume support: Yes
|
- 断点续传: 是
|
||||||
|
|||||||
@@ -1,71 +1,91 @@
|
|||||||
---
|
---
|
||||||
description: Summarize deep research results into markdown report, cover all fields, skip uncertain values.
|
description: 将deep调研结果汇总为markdown报告,覆盖所有字段,跳过不确定值。
|
||||||
allowed-tools: Read, Write, Glob, Bash
|
allowed-tools: Read, Write, Glob, Bash
|
||||||
---
|
---
|
||||||
|
|
||||||
# Research Report - Summary Report
|
# Research Report - 汇总报告
|
||||||
|
|
||||||
## Trigger
|
## 触发方式
|
||||||
`/research-report`
|
`/research-report`
|
||||||
|
|
||||||
## Workflow
|
## 执行流程
|
||||||
|
|
||||||
### Step 1: Locate Results Directory
|
### Step 1: 定位结果目录
|
||||||
Find `*/outline.yaml` in current working directory, read topic and output_dir config.
|
在当前工作目录查找 `*/outline.yaml`,读取topic和output_dir配置。
|
||||||
|
|
||||||
### Step 2: Generate Python Conversion Script
|
### Step 2: 扫描可选摘要字段
|
||||||
Generate `generate_report.py` in `{topic}/` directory, script requirements:
|
读取所有JSON结果,提取适合在目录中显示的字段(数值型、简短指标),例如:
|
||||||
- Read all JSON from output_dir
|
- github_stars
|
||||||
- Read fields.yaml to get field structure
|
- google_scholar_cites
|
||||||
- Cover all field values from each JSON
|
- swe_bench_score
|
||||||
- Skip fields with values containing [uncertain]
|
- user_scale
|
||||||
- Skip fields listed in uncertain array
|
- valuation
|
||||||
- Generate markdown report format: Table of contents (with anchor links) + Detailed content (by field category)
|
- release_date
|
||||||
- Save to `{topic}/report.md`
|
|
||||||
|
|
||||||
#### Script Technical Requirements (Must Follow)
|
使用AskUserQuestion询问用户:
|
||||||
|
- 目录中除了item名称外,还需要显示哪些字段?
|
||||||
|
- 提供动态选项列表(基于实际JSON中存在的字段)
|
||||||
|
|
||||||
**1. JSON Structure Compatibility**
|
### Step 3: 生成Python转换脚本
|
||||||
Support two JSON structures:
|
在 `{topic}/` 目录下生成 `generate_report.py`,脚本要求:
|
||||||
- Flat structure: Fields directly at top level `{"name": "xxx", "release_date": "xxx"}`
|
- 读取output_dir下所有JSON
|
||||||
- Nested structure: Fields in category sub-dict `{"basic_info": {"name": "xxx"}, "technical_features": {...}}`
|
- 读取fields.yaml获取字段结构
|
||||||
|
- 覆盖每个JSON的所有字段值
|
||||||
|
- 跳过值包含[不确定]的字段
|
||||||
|
- 跳过uncertain数组中列出的字段
|
||||||
|
- 生成markdown报告格式:目录(带锚点跳转+用户选择的摘要字段)+ 详细内容(按字段分类)
|
||||||
|
- 保存到 `{topic}/report.md`
|
||||||
|
|
||||||
Field lookup order: Top level -> category mapping key -> Traverse all nested dicts
|
**目录格式要求**:
|
||||||
|
- 必须包含每一个item
|
||||||
|
- 每个item显示:序号、名称(锚点链接)、用户选择的摘要字段
|
||||||
|
- 示例:`1. [GitHub Copilot](#github-copilot) - Stars: 10k | Score: 85%`
|
||||||
|
|
||||||
**2. Category Multi-language Mapping**
|
#### 脚本技术要点(必须遵循)
|
||||||
fields.yaml category names and JSON keys can be any combination (CN-CN, CN-EN, EN-CN, EN-EN). Must establish bidirectional mapping:
|
|
||||||
|
**1. JSON结构兼容**
|
||||||
|
支持两种JSON结构:
|
||||||
|
- 扁平结构:字段直接在顶层 `{"name": "xxx", "release_date": "xxx"}`
|
||||||
|
- 嵌套结构:字段在category子dict中 `{"basic_info": {"name": "xxx"}, "technical_features": {...}}`
|
||||||
|
|
||||||
|
字段查找顺序:顶层 -> category映射key -> 遍历所有嵌套dict
|
||||||
|
|
||||||
|
**2. Category多语言映射**
|
||||||
|
fields.yaml的category名与JSON的key可能是任意组合(中中、中英、英中、英英)。必须建立双向映射:
|
||||||
```python
|
```python
|
||||||
CATEGORY_MAPPING = {
|
CATEGORY_MAPPING = {
|
||||||
"Basic Info": ["basic_info", "Basic Info"],
|
"基本信息": ["basic_info", "基本信息"],
|
||||||
"Technical Features": ["technical_features", "technical_characteristics", "Technical Features"],
|
"技术特性": ["technical_features", "technical_characteristics", "技术特性"],
|
||||||
"Performance Metrics": ["performance_metrics", "performance", "Performance Metrics"],
|
"性能指标": ["performance_metrics", "performance", "性能指标"],
|
||||||
"Milestone Significance": ["milestone_significance", "milestones", "Milestone Significance"],
|
"里程碑意义": ["milestone_significance", "milestones", "里程碑意义"],
|
||||||
"Business Info": ["business_info", "commercial_info", "Business Info"],
|
"商业信息": ["business_info", "commercial_info", "商业信息"],
|
||||||
"Competition & Ecosystem": ["competition_ecosystem", "competition", "Competition & Ecosystem"],
|
"竞争与生态": ["competition_ecosystem", "competition", "竞争与生态"],
|
||||||
"History": ["history", "History"],
|
"历史沿革": ["history", "历史沿革"],
|
||||||
"Market Positioning": ["market_positioning", "market", "Market Positioning"],
|
"市场定位": ["market_positioning", "market", "市场定位"],
|
||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
**3. Complex Value Formatting**
|
**3. 复杂值格式化**
|
||||||
- list of dicts (e.g., key_events, funding_history): Format each dict as one line, separate kv with ` | `
|
- list of dicts(如key_events, funding_history):每个dict格式化为一行,用` | `分隔kv
|
||||||
- Normal list: Short lists joined with comma, long lists displayed with line breaks
|
- 普通list:短列表用逗号连接,长列表换行显示
|
||||||
- Nested dict: Recursive formatting, display with semicolon or line breaks
|
- 嵌套dict:递归格式化,用分号或换行显示
|
||||||
|
- 长文本字符串(超过100字符):添加换行符`<br>`或使用blockquote格式,提高可读性
|
||||||
|
|
||||||
**4. Extra Fields Collection**
|
**4. 额外字段收集**
|
||||||
Collect fields that exist in JSON but not defined in fields.yaml, put in "Other Info" category. Note to filter:
|
收集JSON中有但fields.yaml中没定义的字段,放入"其他信息"分类。注意过滤:
|
||||||
- Internal fields: `_source_file`, `uncertain`
|
- 内部字段:`_source_file`, `uncertain`
|
||||||
- Nested structure top-level keys: `basic_info`, `technical_features` etc.
|
- 嵌套结构顶级key:`basic_info`, `technical_features`等
|
||||||
|
- `uncertain_fields`列表:需要逐行显示每个字段名,不要压缩成一行
|
||||||
|
|
||||||
**5. Uncertain Value Skipping**
|
**5. 不确定值跳过**
|
||||||
Skip conditions:
|
跳过条件:
|
||||||
- Field value contains `[uncertain]` string
|
- 字段值包含`[不确定]`字符串
|
||||||
- Field name is in `uncertain` array
|
- 字段名在`uncertain`数组中
|
||||||
- Field value is None or empty string
|
- 字段值为None或空字符串
|
||||||
|
|
||||||
### Step 3: Execute Script
|
### Step 4: 执行脚本
|
||||||
Run `python {topic}/generate_report.py`
|
运行 `python {topic}/generate_report.py`
|
||||||
|
|
||||||
## Output
|
## 输出
|
||||||
- `{topic}/generate_report.py` - Conversion script
|
- `{topic}/generate_report.py` - 转换脚本
|
||||||
- `{topic}/report.md` - Summary report
|
- `{topic}/report.md` - 汇总报告
|
||||||
|
|||||||
@@ -1,8 +1,8 @@
|
|||||||
#!/usr/bin/env python3
|
#!/usr/bin/env python3
|
||||||
# -*- coding: utf-8 -*-
|
# -*- coding: utf-8 -*-
|
||||||
"""
|
"""
|
||||||
JSON Field Validation Script
|
JSON字段验证脚本
|
||||||
Validate if JSON files completely cover all fields defined in fields.yaml
|
验证JSON文件是否完整覆盖fields.yaml中定义的所有字段
|
||||||
"""
|
"""
|
||||||
|
|
||||||
import json
|
import json
|
||||||
@@ -11,25 +11,25 @@ import sys
|
|||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
from typing import Dict, List, Set, Any, Tuple
|
from typing import Dict, List, Set, Any, Tuple
|
||||||
|
|
||||||
# Category mapping (supports both Chinese and English)
|
# Category中英文映射
|
||||||
CATEGORY_MAPPING = {
|
CATEGORY_MAPPING = {
|
||||||
"Basic Info": ["basic_info", "Basic Info"],
|
"基本信息": ["basic_info", "基本信息"],
|
||||||
"Technical Features": ["technical_features", "technical_characteristics", "Technical Features"],
|
"技术特性": ["technical_features", "technical_characteristics", "技术特性"],
|
||||||
"Performance Metrics": ["performance_metrics", "performance", "Performance Metrics"],
|
"性能指标": ["performance_metrics", "performance", "性能指标"],
|
||||||
"Milestone Significance": ["milestone_significance", "milestones", "Milestone Significance"],
|
"里程碑意义": ["milestone_significance", "milestones", "里程碑意义"],
|
||||||
"Business Info": ["business_info", "commercial_info", "Business Info"],
|
"商业信息": ["business_info", "commercial_info", "商业信息"],
|
||||||
"Competition & Ecosystem": ["competition_ecosystem", "competition", "Competition & Ecosystem"],
|
"竞争与生态": ["competition_ecosystem", "competition", "竞争与生态"],
|
||||||
"History": ["history", "History"],
|
"历史沿革": ["history", "历史沿革"],
|
||||||
"Market Positioning": ["market_positioning", "market", "Market Positioning"],
|
"市场定位": ["market_positioning", "market", "市场定位"],
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
def load_fields_yaml(fields_path: Path) -> Tuple[Set[str], Set[str], Dict[str, str]]:
|
def load_fields_yaml(fields_path: Path) -> Tuple[Set[str], Set[str], Dict[str, str]]:
|
||||||
"""
|
"""
|
||||||
Load fields.yaml, returns:
|
加载fields.yaml,返回:
|
||||||
- all_fields: Set of all field names
|
- all_fields: 所有字段名集合
|
||||||
- required_fields: Set of field names where required=true
|
- required_fields: required=true的字段名集合
|
||||||
- field_categories: Mapping from field name to category
|
- field_categories: 字段名到类别的映射
|
||||||
"""
|
"""
|
||||||
with open(fields_path, 'r', encoding='utf-8') as f:
|
with open(fields_path, 'r', encoding='utf-8') as f:
|
||||||
data = yaml.safe_load(f)
|
data = yaml.safe_load(f)
|
||||||
@@ -52,12 +52,12 @@ def load_fields_yaml(fields_path: Path) -> Tuple[Set[str], Set[str], Dict[str, s
|
|||||||
|
|
||||||
def extract_json_fields(data: Dict, category_mapping: Dict = None) -> Set[str]:
|
def extract_json_fields(data: Dict, category_mapping: Dict = None) -> Set[str]:
|
||||||
"""
|
"""
|
||||||
Extract all field names from JSON (supports flat and nested structures)
|
从JSON中提取所有字段名(支持扁平和嵌套结构)
|
||||||
"""
|
"""
|
||||||
if category_mapping is None:
|
if category_mapping is None:
|
||||||
category_mapping = CATEGORY_MAPPING
|
category_mapping = CATEGORY_MAPPING
|
||||||
|
|
||||||
# Get all possible nested keys
|
# 获取所有可能的嵌套key
|
||||||
nested_keys = set()
|
nested_keys = set()
|
||||||
for keys in category_mapping.values():
|
for keys in category_mapping.values():
|
||||||
nested_keys.update(keys)
|
nested_keys.update(keys)
|
||||||
@@ -66,10 +66,10 @@ def extract_json_fields(data: Dict, category_mapping: Dict = None) -> Set[str]:
|
|||||||
|
|
||||||
def collect_fields(d: Dict, is_top_level: bool = True):
|
def collect_fields(d: Dict, is_top_level: bool = True):
|
||||||
for k, v in d.items():
|
for k, v in d.items():
|
||||||
# Skip internal fields
|
# 跳过内部字段
|
||||||
if k in {"_source_file", "uncertain"}:
|
if k in {"_source_file", "uncertain"}:
|
||||||
continue
|
continue
|
||||||
# If it's a nested structure top-level key, recurse into it
|
# 如果是嵌套结构的顶级key,递归进入
|
||||||
if is_top_level and k in nested_keys:
|
if is_top_level and k in nested_keys:
|
||||||
if isinstance(v, dict):
|
if isinstance(v, dict):
|
||||||
collect_fields(v, is_top_level=False)
|
collect_fields(v, is_top_level=False)
|
||||||
@@ -85,27 +85,27 @@ def extract_json_fields(data: Dict, category_mapping: Dict = None) -> Set[str]:
|
|||||||
def validate_json(json_path: Path, all_fields: Set[str], required_fields: Set[str],
|
def validate_json(json_path: Path, all_fields: Set[str], required_fields: Set[str],
|
||||||
field_categories: Dict[str, str]) -> Dict:
|
field_categories: Dict[str, str]) -> Dict:
|
||||||
"""
|
"""
|
||||||
Validate a single JSON file
|
验证单个JSON文件
|
||||||
Returns validation result dictionary
|
返回验证结果字典
|
||||||
"""
|
"""
|
||||||
with open(json_path, 'r', encoding='utf-8') as f:
|
with open(json_path, 'r', encoding='utf-8') as f:
|
||||||
data = json.load(f)
|
data = json.load(f)
|
||||||
|
|
||||||
json_fields = extract_json_fields(data)
|
json_fields = extract_json_fields(data)
|
||||||
|
|
||||||
# Calculate coverage
|
# 计算覆盖情况
|
||||||
covered = all_fields & json_fields
|
covered = all_fields & json_fields
|
||||||
missing = all_fields - json_fields
|
missing = all_fields - json_fields
|
||||||
extra = json_fields - all_fields
|
extra = json_fields - all_fields
|
||||||
|
|
||||||
# Categorize missing fields
|
# 分类缺失字段
|
||||||
missing_required = missing & required_fields
|
missing_required = missing & required_fields
|
||||||
missing_optional = missing - required_fields
|
missing_optional = missing - required_fields
|
||||||
|
|
||||||
# Group missing fields by category
|
# 按类别分组缺失字段
|
||||||
missing_by_category = {}
|
missing_by_category = {}
|
||||||
for field in missing:
|
for field in missing:
|
||||||
cat = field_categories.get(field, "Unknown")
|
cat = field_categories.get(field, "未知")
|
||||||
if cat not in missing_by_category:
|
if cat not in missing_by_category:
|
||||||
missing_by_category[cat] = []
|
missing_by_category[cat] = []
|
||||||
missing_by_category[cat].append(field)
|
missing_by_category[cat].append(field)
|
||||||
@@ -121,65 +121,65 @@ def validate_json(json_path: Path, all_fields: Set[str], required_fields: Set[st
|
|||||||
"missing_optional": list(missing_optional),
|
"missing_optional": list(missing_optional),
|
||||||
"missing_by_category": missing_by_category,
|
"missing_by_category": missing_by_category,
|
||||||
"extra_fields": list(extra),
|
"extra_fields": list(extra),
|
||||||
"valid": len(missing_required) == 0, # Valid if all required fields are covered
|
"valid": len(missing_required) == 0, # required字段全覆盖则valid
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
def print_result(result: Dict, verbose: bool = True):
|
def print_result(result: Dict, verbose: bool = True):
|
||||||
"""Print validation result"""
|
"""打印验证结果"""
|
||||||
status = "PASS" if result["valid"] else "FAIL"
|
status = "PASS" if result["valid"] else "FAIL"
|
||||||
print(f"\n{'='*60}")
|
print(f"\n{'='*60}")
|
||||||
print(f"[{status}] {result['file']}")
|
print(f"[{status}] {result['file']}")
|
||||||
print(f"{'='*60}")
|
print(f"{'='*60}")
|
||||||
print(f"Coverage: {result['coverage_rate']:.1f}% ({result['covered']}/{result['total_defined']})")
|
print(f"覆盖率: {result['coverage_rate']:.1f}% ({result['covered']}/{result['total_defined']})")
|
||||||
|
|
||||||
if result["missing_required"]:
|
if result["missing_required"]:
|
||||||
print(f"\n[ERROR] Missing required fields ({len(result['missing_required'])}):")
|
print(f"\n[ERROR] 缺失必需字段 ({len(result['missing_required'])}):")
|
||||||
for field in result["missing_required"]:
|
for field in result["missing_required"]:
|
||||||
print(f" - {field}")
|
print(f" - {field}")
|
||||||
|
|
||||||
if verbose and result["missing_optional"]:
|
if verbose and result["missing_optional"]:
|
||||||
print(f"\n[WARN] Missing optional fields ({len(result['missing_optional'])}):")
|
print(f"\n[WARN] 缺失可选字段 ({len(result['missing_optional'])}):")
|
||||||
for cat, fields in result["missing_by_category"].items():
|
for cat, fields in result["missing_by_category"].items():
|
||||||
optional_fields = [f for f in fields if f not in result["missing_required"]]
|
optional_fields = [f for f in fields if f not in result["missing_required"]]
|
||||||
if optional_fields:
|
if optional_fields:
|
||||||
print(f" [{cat}]: {', '.join(optional_fields)}")
|
print(f" [{cat}]: {', '.join(optional_fields)}")
|
||||||
|
|
||||||
if verbose and result["extra_fields"]:
|
if verbose and result["extra_fields"]:
|
||||||
print(f"\n[INFO] Extra fields ({len(result['extra_fields'])}):")
|
print(f"\n[INFO] 额外字段 ({len(result['extra_fields'])}):")
|
||||||
print(f" {', '.join(result['extra_fields'][:10])}")
|
print(f" {', '.join(result['extra_fields'][:10])}")
|
||||||
if len(result["extra_fields"]) > 10:
|
if len(result["extra_fields"]) > 10:
|
||||||
print(f" ... and {len(result['extra_fields']) - 10} more fields")
|
print(f" ... 及 {len(result['extra_fields']) - 10} 个其他字段")
|
||||||
|
|
||||||
|
|
||||||
def main():
|
def main():
|
||||||
"""Main function"""
|
"""主函数"""
|
||||||
import argparse
|
import argparse
|
||||||
parser = argparse.ArgumentParser(description="Validate if JSON files cover all fields defined in fields.yaml")
|
parser = argparse.ArgumentParser(description="验证JSON文件是否覆盖fields.yaml定义的所有字段")
|
||||||
parser.add_argument("--fields", "-f", type=str, help="Path to fields.yaml", default="fields.yaml")
|
parser.add_argument("--fields", "-f", type=str, help="fields.yaml路径", default="fields.yaml")
|
||||||
parser.add_argument("--json", "-j", type=str, nargs="*", help="JSON file paths to validate")
|
parser.add_argument("--json", "-j", type=str, nargs="*", help="要验证的JSON文件路径")
|
||||||
parser.add_argument("--dir", "-d", type=str, help="JSON files directory", default="results")
|
parser.add_argument("--dir", "-d", type=str, help="JSON文件目录", default="results")
|
||||||
parser.add_argument("--quiet", "-q", action="store_true", help="Show summary only")
|
parser.add_argument("--quiet", "-q", action="store_true", help="只显示摘要")
|
||||||
args = parser.parse_args()
|
args = parser.parse_args()
|
||||||
|
|
||||||
# Locate fields.yaml
|
# 定位fields.yaml
|
||||||
fields_path = Path(args.fields)
|
fields_path = Path(args.fields)
|
||||||
if not fields_path.exists():
|
if not fields_path.exists():
|
||||||
# Try to find in current and parent directory
|
# 尝试在当前目录和父目录查找
|
||||||
for p in [Path.cwd() / "fields.yaml", Path.cwd().parent / "fields.yaml"]:
|
for p in [Path.cwd() / "fields.yaml", Path.cwd().parent / "fields.yaml"]:
|
||||||
if p.exists():
|
if p.exists():
|
||||||
fields_path = p
|
fields_path = p
|
||||||
break
|
break
|
||||||
|
|
||||||
if not fields_path.exists():
|
if not fields_path.exists():
|
||||||
print(f"[ERROR] fields.yaml not found: {fields_path}")
|
print(f"[ERROR] fields.yaml不存在: {fields_path}")
|
||||||
sys.exit(1)
|
sys.exit(1)
|
||||||
|
|
||||||
print(f"Field definitions file: {fields_path}")
|
print(f"字段定义文件: {fields_path}")
|
||||||
all_fields, required_fields, field_categories = load_fields_yaml(fields_path)
|
all_fields, required_fields, field_categories = load_fields_yaml(fields_path)
|
||||||
print(f"Total fields: {len(all_fields)} (required: {len(required_fields)}, optional: {len(all_fields) - len(required_fields)})")
|
print(f"总字段数: {len(all_fields)} (必需: {len(required_fields)}, 可选: {len(all_fields) - len(required_fields)})")
|
||||||
|
|
||||||
# Collect JSON files
|
# 收集JSON文件
|
||||||
json_files = []
|
json_files = []
|
||||||
if args.json:
|
if args.json:
|
||||||
json_files = [Path(p) for p in args.json]
|
json_files = [Path(p) for p in args.json]
|
||||||
@@ -189,27 +189,27 @@ def main():
|
|||||||
json_files = sorted(json_dir.glob("*.json"))
|
json_files = sorted(json_dir.glob("*.json"))
|
||||||
|
|
||||||
if not json_files:
|
if not json_files:
|
||||||
print(f"[WARN] No JSON files found")
|
print(f"[WARN] 未找到JSON文件")
|
||||||
sys.exit(0)
|
sys.exit(0)
|
||||||
|
|
||||||
# Validate each file
|
# 验证每个文件
|
||||||
results = []
|
results = []
|
||||||
for json_path in json_files:
|
for json_path in json_files:
|
||||||
if not json_path.exists():
|
if not json_path.exists():
|
||||||
print(f"[WARN] File not found: {json_path}")
|
print(f"[WARN] 文件不存在: {json_path}")
|
||||||
continue
|
continue
|
||||||
result = validate_json(json_path, all_fields, required_fields, field_categories)
|
result = validate_json(json_path, all_fields, required_fields, field_categories)
|
||||||
results.append(result)
|
results.append(result)
|
||||||
print_result(result, verbose=not args.quiet)
|
print_result(result, verbose=not args.quiet)
|
||||||
|
|
||||||
# Summary
|
# 汇总
|
||||||
print(f"\n{'='*60}")
|
print(f"\n{'='*60}")
|
||||||
print("Summary")
|
print("汇总")
|
||||||
print(f"{'='*60}")
|
print(f"{'='*60}")
|
||||||
passed = sum(1 for r in results if r["valid"])
|
passed = sum(1 for r in results if r["valid"])
|
||||||
avg_coverage = sum(r["coverage_rate"] for r in results) / len(results) if results else 0
|
avg_coverage = sum(r["coverage_rate"] for r in results) / len(results) if results else 0
|
||||||
print(f"Validation passed: {passed}/{len(results)}")
|
print(f"验证通过: {passed}/{len(results)}")
|
||||||
print(f"Average coverage: {avg_coverage:.1f}%")
|
print(f"平均覆盖率: {avg_coverage:.1f}%")
|
||||||
|
|
||||||
if passed < len(results):
|
if passed < len(results):
|
||||||
sys.exit(1)
|
sys.exit(1)
|
||||||
|
|||||||
Reference in New Issue
Block a user