走完初步流程

This commit is contained in:
2026-06-06 07:44:26 +08:00
parent 580fe4ad6e
commit e0d6bdc9c3
51 changed files with 6883 additions and 2410 deletions

462
example-analysis.md Normal file
View File

@@ -0,0 +1,462 @@
# example.ts 代码分析文档
## 一、文件概述
`example.ts` 是 LLMWiki 项目的**核心导入管道Ingest Pipeline**实现负责将源文档PDF/DOCX/PPTX/Markdown 等)自动转换为结构化的 Wiki 页面。
**核心职责:**
- 读取源文件内容
- 提取嵌入图片并生成描述
- 通过 LLM 两阶段分析(分析 → 生成)将源文档转化为 Wiki 页面
- 处理长文档的分块分析
- 写入文件并管理缓存
- 生成嵌入向量(可选)
---
## 二、整体架构
```
源文件输入
┌─────────────────────────────────────────────────┐
│ autoIngest() │
│ (项目级锁入口) │
│ │
│ ┌─ Step 0: 缓存检查 │
│ │ ├─ HIT → 跳至图片处理 → 返回 │
│ │ └─ MISS → 继续全量管道 │
│ │ │
│ ┌─ Step 0.5: 图片提取 │
│ │ └─ extractAndSaveSourceImages() │
│ │ │
│ ┌─ Step 0.6: 图片描述Caption
│ │ └─ captionMarkdownImages() │
│ │ │
│ ┌─ 长文档预算检查 │
│ │ ├─ 超长 → analyzeLongSourceInChunks() │
│ │ └─ 正常 → 直接传入 │
│ │ │
│ ┌─ Step 1: 分析阶段 │
│ │ └─ streamChat(buildAnalysisPrompt()) │
│ │ │
│ ┌─ Step 2: 生成阶段 │
│ │ └─ streamChat(buildGenerationPrompt()) │
│ │ │
│ ┌─ Step 2.5: Review 建议(可选) │
│ │ └─ streamChat(buildReviewSuggestionPrompt()) │
│ │ │
│ ┌─ Step 3: 文件写入 │
│ │ └─ writeFileBlocks() │
│ │ │
│ ┌─ Step 3.5: 图片注入源摘要页 │
│ │ └─ injectImagesIntoSourceSummary() │
│ │ │
│ ┌─ Step 4: Review 项解析 │
│ │ └─ parseReviewBlocks() │
│ │ │
│ ┌─ Step 5: 缓存保存 │
│ │ └─ saveIngestCache() │
│ │ │
│ ┌─ Step 6: 嵌入生成(可选) │
│ │ └─ embedPage() │
│ └───────────────────────────────────────────────────┘
Wiki 页面输出
```
---
## 三、核心流程详解
### 3.1 入口:`autoIngest()`
```typescript
autoIngest(
projectPath: string, // 项目根目录
sourcePath: string, // 源文件路径
llmConfig: LlmConfig, // LLM 配置
signal?: AbortSignal, // 取消信号
folderContext?: string, // 文件夹上下文(用于分类提示)
): Promise<string[]> // 返回已写入文件路径列表
```
**关键设计:**
- 使用 `withProjectLock()` 确保同一项目不会并发执行多个导入任务
- 锁的必要性:分析阶段读取 `wiki/index.md`,生成阶段覆盖它,不串行化会导致互相覆盖
### 3.2 缓存机制
**检查逻辑:**
```
checkIngestCache(projectPath, sourceIdentity, sourceContent)
├─ 源内容未变化 → 返回缓存的文件列表(跳过 LLM 调用)
└─ 源内容有变化 → 返回 null走全量管道
```
**缓存命中时的行为:**
- 仍然执行图片提取(用户可能在旧版本导入过,当时没有图片功能)
- 仍然执行图片描述注入
- 跳过 Step 1~2LLM 分析 + 生成)
### 3.3 图片处理管道
#### Step 0.5:图片提取
```
extractAndSaveSourceImages(projectPath, sourcePath, sourceSummarySlug)
wiki/media/<source-slug>/image1.png
wiki/media/<source-slug>/image2.png
...
```
**支持格式:** PDF、PPTX、DOCX通过 pdfium 等后端提取)
**去重策略:**
- 使用 `ingestImageExtractionPromises` Map 缓存 Promise
- 键 = `项目路径 + 源路径 + slug + 文件大小:修改时间`
- 最大 32 个缓存条目,超出时淘汰最旧
#### Step 0.6图片描述Caption
```
captionMarkdownImages(projectPath, enrichedSourceContent, captionLlm, {
shouldCaption: (url) => url.startsWith(ourMediaPrefix), // 只处理当前导入的图片
urlToAbsPath: (url) => promptImageUrlToAbs(pp, url),
concurrency: mmCfg.concurrency,
onProgress: (done, total) => activity.updateItem(...),
})
```
**为什么需要描述:**
-`alt` 的图片在文本摘要时会被 LLM 静默丢弃
- 有了描述后alt 文本携带足够语义,生成 LLM 会保留图片引用
**描述缓存:**
- SHA-256 键值缓存(`.llm-wiki/image-caption-cache.json`
- 跨文档去重(共享的 logo/图表模板只描述一次)
**主开关行为:**
-`multimodalConfig.enabled = false` 时:
- 不仅跳过 Caption LLM 调用
- 还会从 `sourceContent` 中剥离 `![](url)` 引用
- 跳过后续的图片注入
- 净效果Wiki 侧完全不引用图片
### 3.4 长文档处理
**触发条件:** `enrichedSourceContent.length > sourceBudget`
**预算计算:**
```typescript
sourceBudget = maxCtx - responseReserve - stableReserve - instructionReserve
// 范围限制在 [8000, min(300000, maxCtx * 0.6)]
```
**分块策略:**
```
splitSourceIntoSemanticChunks(content, targetChars, overlapChars)
按段落/章节边界分割,每块带 8% 重叠800~3000 字符)
```
**分块分析流程:**
```
对每个 chunk
├─ streamChat(buildChunkAnalysisSystemPrompt(), buildChunkAnalysisUserPrompt())
├─ 提取 "Chunk Analysis" 和 "Updated Global Digest"
├─ 更新全局摘要(最长 15000 字符)
└─ 保存检查点(支持中断恢复)
```
**检查点机制:**
- 路径:`.llm-wiki/ingest-progress/<slug>-<hash>.json`
- 包含:源哈希、块数量、已完成进度、全局摘要、各块分析
- 兼容性验证:版本、源哈希、源长度、块参数全部匹配才恢复
**最终合并:**
```
Consolidated Long-Document Analysis
├─ Final Global Digest
└─ Per-Chunk Analyses
```
### 3.5 Step 1分析阶段
**System Prompt 结构:**
```
buildAnalysisPrompt(purpose, index, sourceContent)
├─ 角色设定:专业研究分析师
├─ 语言规则
├─ 分析框架:
│ ├─ 关键实体(人物、组织、产品、数据集、工具)
│ ├─ 关键概念(理论、方法、技术、现象)
│ ├─ 主要论点与发现
│ ├─ 与现有 Wiki 的关联
│ ├─ 矛盾与张力
│ └─ 建议
├─ Wiki 目的(可选)
└─ 当前 Wiki 索引(可选)
```
**User Message 结构:**
```
Analyze this source document:
**File:** ${sourceIdentity}
**Folder context:** ${folderContext} (可选)
---
${sourceContext}
```
**LLM 参数:**
- `temperature: 0.1`(低随机性,确保一致性)
- `reasoning.mode: "off"`(关闭推理模式)
- `max_tokens: 4096`
### 3.6 Step 2生成阶段
**System Prompt 结构:**
```
buildGenerationPrompt(schema, purpose, index, sourceFileName, ...)
├─ 角色设定Wiki 维护者
├─ 语言规则
├─ 源文件信息
├─ 项目 Schema如有
├─ 生成清单:
│ 1. 源摘要页wiki/sources/<slug>.md
│ 2. 实体页wiki/entities/ 或 schema 定义目录)
│ 3. 概念页wiki/concepts/ 或 schema 定义目录)
│ 4. 更新 wiki/index.md
│ 5. 日志条目wiki/log.md
│ 6. 更新 wiki/overview.md
├─ Frontmatter 规则(严格 YAML 格式)
├─ Review 块类型
└─ 输出格式FILE 块 + 可选 REVIEW 块)
```
**User Message 结构:**
```
Source document to process: **${sourceIdentity}**
The Stage 1 analysis below is CONTEXT to inform your output.
Do NOT echo its tables, bullet points, or prose.
## Stage 1 Analysis (context only — do not repeat)
${analysis}
## Source Context
${sourceContext}
---
Now emit the FILE blocks for the wiki files derived from **${sourceIdentity}**.
Your response MUST begin with `---FILE:` as the very first characters.
```
**LLM 参数:**
- `max_tokens` 根据上下文窗口动态计算:
- 默认8192
- 128K 上下文16384
- 256K 上下文24576
- 512K 上下文32768
### 3.7 Step 2.5Review 建议(可选)
**触发条件:**
```typescript
shouldRunDedicatedReviewStage(generation):
generation.length >= 10000 // 响应足够长
|| countFileBlocks(generation) >= 4 // 生成了足够多文件
|| /---REVIEW:/.test(generation) // 已有 REVIEW 块
```
**System Prompt**
```
buildReviewSuggestionPrompt(purpose, index, sourceIdentity, analysis, sourceContext, generation)
├─ 角色:识别高价值后续研究项
├─ 仅输出 REVIEW 块(不生成 Wiki 页面)
├─ 来源上下文(截断到预算内)
└─ 生成的 Wiki 输出(截断到预算内)
```
### 3.8 Step 3文件写入
**解析 FILE 块:**
```typescript
parseFileBlocks(text): ParseFileBlocksResult
blocks: ParsedFileBlock[]
warnings: string[]
```
**解析器修复的问题:**
| 问题 | 描述 | 修复方式 |
|------|------|----------|
| H1 | Windows CRLF 行尾 | 先规范化为 LF |
| H2 | 流截断导致最后一个块丢失 | 发出警告 |
| H3 | 标记空白/大小写变体 | 正则不区分大小写,容忍空白 |
| H5 | 代码块内的字面量 `---END FILE---` | 追踪 fence 状态fence 内不关闭 |
| H6 | 空路径块静默丢弃 | 发出警告 |
**路径安全校验:**
```typescript
isSafeIngestPath(p): boolean
wiki/
..
Windows
CONPRNAUXNULCOM1-9LPT1-9
NUL/
```
**写入策略(按路径类型):**
| 路径类型 | 策略 |
|----------|------|
| `wiki/log.md` | 追加模式append |
| `wiki/index.md` / `wiki/overview.md` | 全覆盖overwrite |
| 内容页entities/concepts/等) | 合并模式merge |
**合并策略三层:**
1. Frontmatter 数组字段sources、tags、related→ 并集合并
2. 正文内容不同 → LLM 生成合并版本
3. 锁定字段type、title、created→ 保留原值updated 更新为当天
**合并失败回退:**
- LLM 合并失败 → 使用传入正文 + 数组字段并集
- 合并前备份原文件到 `.llm-wiki/page-history/`
**语言守卫:**
```typescript
contentMatchesTargetLanguage(content, target): boolean
frontmatter + +
CJK CJK Latin Latin
```
### 3.9 Step 4~6收尾
**Step 4Review 项解析**
```typescript
parseReviewBlocks(text, sourcePath): ReviewItem[]
contradiction / duplicate / missing-page / suggestion
OPTIONS
PAGES
SEARCH Deep Research
```
**Step 5缓存保存**
- 仅当无硬失败(磁盘满、权限错误等)时保存
- 软丢弃(语言不匹配、路径穿越拒绝)不影响缓存
**Step 6嵌入生成**
- 跳过 index、log、overview 页面
- 从 frontmatter 提取 title
- 调用 `embedPage()` 生成向量
---
## 四、辅助入口:`startIngest()` 和 `executeIngestWrites()`
这两个函数服务于**交互式聊天模式**(用户在 UI 中手动触发导入):
### `startIngest()`
- 设置聊天模式为 "ingest"
- 提前提取图片(与 LLM 分析并行)
- 发送分析请求并流式展示结果
### `executeIngestWrites()`
- 接收聊天历史作为上下文
- 构建生成提示
- 解析 FILE 块并写入文件
- 注入图片引用到源摘要页
---
## 五、关键数据结构
### FILE 块格式
```
---FILE: wiki/path/to/page.md---
---
type: entity
title: Example Entity
created: 2026-04-29
updated: 2026-04-29
tags: [example, demo]
related: [related-slug-1]
sources: ["source-file.pdf"]
---
# Example Entity
Body content here.
---END FILE---
```
### REVIEW 块格式
```
---REVIEW: suggestion | Precise Title---
Description of the gap and why it matters.
OPTIONS: Create Page | Skip
PAGES: wiki/page1.md, wiki/page2.md
SEARCH: query 1 | query 2 | query 3
---END REVIEW---
```
### 检查点结构
```typescript
interface LongSourceCheckpoint {
version: 1
sourceIdentity: string
sourceHash: string // FNV-1a 64-bit hash
sourceLength: number
sourceBudget: number
targetChars: number
overlapChars: number
chunkTotal: number
completedThrough: number // 已完成到第几个块
globalDigest: string // 全局摘要
analyses: string[] // 各块分析
updatedAt: number
}
```
---
## 六、常量配置
| 常量 | 值 | 说明 |
|------|-----|------|
| `LONG_SOURCE_MIN_BUDGET` | 8,000 | 长文档最小预算 |
| `LONG_SOURCE_MAX_SINGLE_PASS_BUDGET` | 300,000 | 长文档单次最大预算 |
| `LONG_SOURCE_CHUNK_MIN` | 12,000 | 分块最小大小 |
| `LONG_SOURCE_CHUNK_MAX` | 60,000 | 分块最大大小 |
| `LONG_SOURCE_DIGEST_MAX` | 15,000 | 全局摘要最大长度 |
| `LONG_SOURCE_CHUNK_ANALYSIS_MAX` | 40,000 | 单块分析最大长度 |
| `REVIEW_STAGE_MIN_SIGNAL_CHARS` | 10,000 | Review 阶段最小信号字符 |
| `REVIEW_STAGE_MIN_FILE_BLOCKS` | 4 | Review 阶段最小文件块数 |
---
## 七、与 Python 项目的映射关系
| TypeScript 模块 | Python 对应模块 | 状态 |
|----------------|----------------|------|
| `autoIngest()` | `src/main.py::main()` | 部分实现 |
| `buildAnalysisPrompt()` | `src/prompt.py::build_analysis_prompt()` | 已实现 |
| `sourceIdentityForPath()` | `src/source.py::source_identity_for_path()` | 已实现 |
| `streamChat()` | `src/llm.py::query_analysis()` | 已实现 |
| `parseFileBlocks()` | 待实现 | 未实现 |
| `writeFileBlocks()` | 待实现 | 未实现 |
| `splitSourceIntoSemanticChunks()` | 待实现 | 未实现 |
| `analyzeLongSourceInChunks()` | 待实现 | 未实现 |
| `checkIngestCache()` | 待实现 | 未实现 |
| 图片提取管道 | 待实现 | 未实现 |