diff --git a/AI_REASONING_CONTINUATION_RESEARCH.md b/AI_REASONING_CONTINUATION_RESEARCH.md
new file mode 100644
index 0000000..a180dd5
--- /dev/null
+++ b/AI_REASONING_CONTINUATION_RESEARCH.md
@@ -0,0 +1,242 @@
+# AI reasoning_content 回传与“继续思考”实验研究报告
+
+日期:2026-08-23
+范围:OpenCodeGo / `deepseek-v4-flash`,OpenAI 兼容 `/chat/completions`
+状态:实验性研究,未集成到主项目
+
+## 摘要
+
+本报告记录了为验证“将模型上一轮 `reasoning_content` 截断/修改后回传,模型是否能够继续合理思考”而执行的一系列实验。
+
+结论是:**仅发送 `reasoning_content` 或“上一轮思维链”并不稳定;成功率最高的方案是伪造一段 tool call 历史,并把续写要求、输出格式和需要引用的内容放入 tool 结果中。**
+
+与本研究相关的代码在 `AiV2.Tests/Program.cs` 中,默认不会执行。
+
+## 目标
+
+验证以下问题:
+
+1. 模型能否在不追加新的 user 消息时,基于 assistant 的 `reasoning_content` 继续生成最终回答?
+2. 截断后的 `reasoning_content` 是否仍然能被模型读取和引用?
+3. 伪造 `tool_calls` + `tool` 结果历史是否比普通多轮对话更有效?
+4. 工具结果中的措辞是否会显著影响模型对“目标思维链”的定位?
+5. 标记在思维链中的位置是否影响模型的可召回性?
+
+## 实验方法与基础设施
+
+所有实验使用:
+
+- 当前应用配置文件中的 OpenCodeGo / `deepseek-v4-flash`
+- `stream=false` 的一次性 OpenAI 兼容请求
+- 初始问题:
+ `A 比 B 高 20%,B 比 C 高 25%,那么 A 比 C 高多少?`
+- 确定性记忆标记:
+ `TOKEN_MARK=731942`
+
+专用测试位于 `AiV2.Tests/Program.cs`,默认跳过。
+
+执行方式:
+
+```powershell
+dotnet build .\AiV2.Tests\AiV2.Tests.csproj /p:AiV2TestsBuilding=true
+$env:ARR_AI_E2E = "1"
+$env:ARR_AI_E2E_TOOL = "1"
+$env:ARR_AI_TOOL_REPORT_PATH = "AI_reasoning_continuation_position_sweep_report.md"
+& .\AiV2.Tests\bin\Debug\net461\AiV2.Tests.exe
+```
+
+## 尝试过程
+
+### 1. 普通多轮:assistant 推理 + 正文再回传
+
+消息结构:
+
+```text
+system → user → assistant(reasoning_content + content) → user
+```
+
+结果:协议层接受 `reasoning_content`,模型通常会重新推导,但无法确认它真正延续了旧思维链。
+
+### 2. 只发 assistant 推理,不追加 user
+
+消息结构:
+
+```text
+system → user → assistant(reasoning_content,content 不发送)
+```
+
+结果:OpenCodeGo 接受最后一条 assistant 消息并生成完成,但新 reasoning 往往从题目重新开始。
+
+### 3. 伪造 tool call 历史
+
+消息结构:
+
+```text
+system
+→ user
+→ assistant(content 不发送, reasoning_content, tool_calls)
+→ tool(tool_call_id, 工具结果)
+```
+
+请求同时声明 `tools`,并将 `tool_choice` 设置为 `"none"`。
+
+工具结果中可以注入:
+
+- “继续当前推理”指令
+- 输出格式要求
+- 需要引用的旧思维链内容
+- “如果没有该内容,明确回答不存在”的边界条件
+
+这一步开始观察到:**模型可能在部分运行中读取并引用 `reasoning_content`。**
+
+### 4. 确定性记忆标记
+
+为避免模型从工具结果中照抄标记,标记值只注入到 `reasoning_content`,`tool` 结果中不出现标记值。
+
+### 5. 措辞 A/B
+
+对比三种措辞:
+
+```text
+A:上一轮思维链
+B:当前推理
+C:调用工具之前的思维链
+D:本次 tool_calls 消息中携带的 reasoning_content
+```
+
+去掉“上一轮”“截断点”等歧义词后,B/D 的命中率明显更高。
+
+### 6. 标记位置扫描
+
+使用成功率最高的两种措辞:
+
+- B:当前推理
+- D:本次 tool_calls 消息中携带的 reasoning_content
+
+标记插入完整思维链的 25%、50%、75%、末尾四个位置,每个组合重复 5 次。
+
+## 结果
+
+### 措辞 A/B(每个组合 3 次)
+
+| 工具提示措辞 | 正确写出标记 | 明确说标记不存在 |
+|---|---:|---:|
+| A:上一轮思维链 | 2 / 3 | 1 / 3 |
+| B:当前推理 | 3 / 3 | 0 / 3 |
+| C:调用工具之前的思维链 | 2 / 3 | 1 / 3 |
+| D:tool_calls 携带的 reasoning_content | 3 / 3 | 0 / 3 |
+
+### 位置扫描(每个组合 5 次)
+
+| 措辞 | 标记位置 | 正确写出标记 | 明确说没有 | 请求失败 |
+|---|---|---:|---:|---:|
+| B | 25% | 4 / 5 | 1 / 5 | 0 |
+| B | 50% | 5 / 5 | 1 / 5 | 0 |
+| B | 75% | 5 / 5 | 1 / 5 | 0 |
+| B | 末尾 | 3 / 5 | 3 / 5 | 0 |
+| D | 25% | 4 / 5 | 0 / 5 | 1(HTTP 503) |
+| D | 50% | 5 / 5 | 0 / 5 | 0 |
+| D | 75% | 5 / 5 | 2 / 5 | 0 |
+| D | 末尾 | 4 / 5 | 2 / 5 | 0 |
+
+## 关键发现
+
+1. **tool call 是当前最有效的载体。**
+
+把“继续推理、输出格式、引用要求”放进 tool 结果,模型会把这些内容当作任务输入,而不是普通 user 消息。
+
+2. **措辞决定模型能否正确定位思维链。**
+
+“当前推理”和“本次 tool_calls 消息中携带的 reasoning_content”比“上一轮思维链”更稳定。不要使用“上一轮”“截断点”这些容易引起歧义的词。
+
+3. **标记在思维链开头/中部时最容易召回。**
+
+50% 和 75% 位置均为 5/5;末尾位置明显下降。这说明不要把关键事实放在 `reasoning_content` 末尾。
+
+4. **“正确写出标记”和“明确说没有”不是互斥的。**
+
+模型有时先否定、后在同一输出中写出标记。统计时两者可能同时为真,人工审阅必须以完整 reasoning 和正文为准。
+
+5. **服务端不稳定。**
+
+D/25% 有一次 HTTP 503,属于服务端错误,不是模型失败。OpenCodeGo 未返回 `usage`,因此无法评估成本/token。
+
+6. **主项目不应直接启用该机制。**
+
+当前实验只能证明“部分情况下有效”,不足以作为生产依赖。
+
+## 推荐方案
+
+### 推荐消息结构
+
+```text
+system
+→ user(原始任务)
+→ assistant(
+ content 不发送,
+ reasoning_content,
+ tool_calls: [{ id, type: "function", function: { name, arguments } }]
+ )
+→ tool(
+ tool_call_id,
+ content: "工具结果:请继续你本次 tool_calls 消息中携带的 reasoning_content。
+ 请原样写出其中的内部记忆标记;如果没有,请明确回答“标记不存在”。
+ 请使用 Markdown 输出。"
+ )
+```
+
+请求级配置:
+
+```json
+{
+ "tools": [
+ {
+ "type": "function",
+ "function": {
+ "name": "analysis_hint",
+ "description": "提供继续上一个内部推理的提示",
+ "parameters": {
+ "type": "object",
+ "properties": {
+ "instruction": { "type": "string" }
+ },
+ "required": ["instruction"]
+ }
+ }
+ }
+ ],
+ "tool_choice": "none"
+}
+```
+
+### 推荐提示词措辞
+
+推荐:
+
+```text
+请继续你当前正在进行的内部推理。
+请继续你本次 tool_calls 消息中携带的 reasoning_content。
+```
+
+不推荐:
+
+```text
+上一轮思维链在这里被截断,请继续。
+```
+
+## 对主项目的建议
+
+- 保留现有 UI 对 `reasoning_content` 的展示。
+- 不要在生产管线中自动发送/截断/修改 `reasoning_content`。
+- 如未来需要接入,优先使用 tool call 载体,并将关键指令、格式和引用要求放入 tool 结果。
+- 不要假设 `reasoning_content` 的末尾内容一定能被模型召回。
+- 增加原始响应/请求日志,以便区分“模型未看到”和“模型看到了但没引用”。
+
+## 专用测试状态
+
+- 测试代码:`AiV2.Tests/Program.cs` 中的 `OpenCodeGoFakeToolCallTests`。
+- 默认行为:不执行。
+- 启用条件:`ARR_AI_E2E=1` 且 `ARR_AI_E2E_TOOL=1`。
+- 输出:Markdown 报告,路径通过 `ARR_AI_TOOL_REPORT_PATH` 指定。
+
+主项目中的实验性 `ChatMessage` 扩展、`AiReasoningContinuationSettings`、UI 开关、回传/降级逻辑和 tool call 历史构造均已移除。
diff --git a/AiV2.Tests/Program.cs b/AiV2.Tests/Program.cs
index b894c71..27a3081 100644
--- a/AiV2.Tests/Program.cs
+++ b/AiV2.Tests/Program.cs
@@ -1,13 +1,20 @@
using System;
using System.Collections.Generic;
using System.Collections.Immutable;
+using System.Diagnostics;
using System.IO;
using System.Linq;
+using System.Net.Http;
+using System.Net.Http.Headers;
using System.Text;
using System.Text.Json;
+using System.Text.Json.Serialization;
+using System.Threading;
+using System.Threading.Tasks;
using AnotherReplayReader;
using AnotherReplayReader.ReplayFile;
using AnotherReplayReader.Utils;
+using AIAnalyze = AiV2.Tests.TestAIAnalyze;
namespace AiV2.Tests
{
@@ -36,6 +43,7 @@ namespace AiV2.Tests
Run("UserKnowledgeOverlay", UserKnowledgeOverlayTests.Run);
Run("PromptBuilders", PromptBuildersTests.Run);
Run("RevisionFactsAndSerialization", RevisionFactsAndSerializationTests.Run);
+ Run("OpenCodeGoFakeToolCallE2e", OpenCodeGoFakeToolCallTests.Run);
Run("UserReplayFactIndexRepro", UserReplayFactIndexReproTests.Run);
Console.WriteLine();
@@ -748,6 +756,662 @@ namespace AiV2.Tests
}
}
+ ///
+ /// AI 分析主工程实验代码已清理;本类只存在于测试工程中。
+ /// 用于直接构造 OpenAI 兼容请求,避免把实验字段带回主项目。
+ ///
+ internal class TestAIAnalyze
+ {
+ private readonly TestAiClient _client = new();
+
+ internal sealed class Result
+ {
+ public Result(string response, string? reasoning)
+ {
+ Response = response;
+ Reasoning = reasoning;
+ }
+
+ public string Response { get; }
+ public string? Reasoning { get; }
+ public int? PromptTokens { get; }
+ public int? CompletionTokens { get; }
+ public int? TotalTokens { get; }
+ public int? ReasoningTokens { get; }
+ }
+
+ internal sealed record ChatMessage(
+ [property: JsonPropertyName("role")] string Role,
+ [property: JsonPropertyName("content"),
+ JsonIgnore(Condition = JsonIgnoreCondition.WhenWritingNull)] string? Content = null,
+ [property: JsonPropertyName("reasoning_content"),
+ JsonIgnore(Condition = JsonIgnoreCondition.WhenWritingNull)] string? ReasoningContent = null,
+ [property: JsonPropertyName("tool_calls"),
+ JsonIgnore(Condition = JsonIgnoreCondition.WhenWritingNull)] IReadOnlyList? ToolCalls = null,
+ [property: JsonPropertyName("tool_call_id"),
+ JsonIgnore(Condition = JsonIgnoreCondition.WhenWritingNull)] string? ToolCallId = null)
+ {
+ internal sealed record ToolCall(
+ [property: JsonPropertyName("id")] string Id,
+ [property: JsonPropertyName("type")] string Type,
+ [property: JsonPropertyName("function")] ToolCallFunction Function);
+
+ internal sealed record ToolCallFunction(
+ [property: JsonPropertyName("name")] string Name,
+ [property: JsonPropertyName("arguments")] string Arguments);
+
+ internal static ChatMessage Assistant(string? content, string? reasoningContent = null) =>
+ new("assistant",
+ string.IsNullOrEmpty(content) ? null : content,
+ string.IsNullOrEmpty(reasoningContent) ? null : reasoningContent);
+
+ internal static ChatMessage AssistantToolCall(
+ string? reasoningContent,
+ string toolCallId,
+ string toolName,
+ string arguments) =>
+ new(
+ "assistant",
+ null,
+ reasoningContent,
+ new[]
+ {
+ new ChatMessage.ToolCall(
+ toolCallId,
+ "function",
+ new ChatMessage.ToolCallFunction(toolName, arguments))
+ },
+ null);
+
+ internal static ChatMessage ToolResult(string toolCallId, string content) =>
+ new("tool", content, null, null, toolCallId);
+ }
+
+ public async Task CompleteAsync(
+ ImmutableList messages,
+ AiRequestContext context,
+ object? unused,
+ CancellationToken cancellationToken)
+ {
+ var response = await _client.CompleteAsync(
+ context.Provider, context.Model, messages, cancellationToken);
+ if (response.Error is not null)
+ {
+ throw new InvalidOperationException(response.Error);
+ }
+ return new Result(response.Content ?? string.Empty, response.Reasoning);
+ }
+
+ public static string BuildOverviewUserPrompt(IReadOnlyList slices) =>
+ AnotherReplayReader.Utils.AIAnalyze.BuildOverviewUserPrompt(slices);
+
+ public static string BuildSegmentUserPromptV2(
+ int segmentIndex,
+ int totalSegments,
+ ReplaySlice slice,
+ int eventCount,
+ string? title,
+ string? description = null,
+ IEnumerable? backqueryHints = null) =>
+ AnotherReplayReader.Utils.AIAnalyze.BuildSegmentUserPromptV2(
+ segmentIndex, totalSegments, slice, eventCount,
+ title, description, backqueryHints);
+
+ public static string BuildSummaryUserPromptV2(int totalEventCount) =>
+ AnotherReplayReader.Utils.AIAnalyze.BuildSummaryUserPromptV2(totalEventCount);
+
+ public static string BuildBackqueryUserPrompt(string sliceText) =>
+ AnotherReplayReader.Utils.AIAnalyze.BuildBackqueryUserPrompt(sliceText);
+
+ public static string BuildRevisionUserPrompt(
+ string draft,
+ string validationIssues,
+ string relevantFacts) =>
+ AnotherReplayReader.Utils.AIAnalyze.BuildRevisionUserPrompt(
+ draft, validationIssues, relevantFacts);
+ }
+
+ internal sealed record TestAiResponse(
+ string? Content,
+ string? Reasoning,
+ string? Error = null,
+ int? StatusCode = null);
+
+ internal sealed class TestAiClient
+ {
+ private readonly HttpClient _http = new()
+ {
+ Timeout = TimeSpan.FromMinutes(5)
+ };
+
+ public async Task CompleteAsync(
+ AiProvider provider,
+ AiModel model,
+ IReadOnlyList messages,
+ CancellationToken cancellationToken)
+ {
+ var request = new Dictionary
+ {
+ ["model"] = model.ModelId,
+ ["temperature"] = provider.DefaultTemperature,
+ ["top_p"] = provider.DefaultTopP,
+ ["max_tokens"] = provider.DefaultMaxTokens,
+ ["stream"] = false
+ };
+ foreach (var kv in model.ExtraParameters)
+ {
+ request[kv.Key] = kv.Value;
+ }
+ request["messages"] = messages.ToArray();
+
+ var uri = new Uri(new(provider.BaseUrl.TrimEnd('/') + "/"), "chat/completions");
+ using var httpRequest = new HttpRequestMessage(HttpMethod.Post, uri);
+ httpRequest.Headers.Authorization = new AuthenticationHeaderValue(
+ "Bearer", provider.ApiKey);
+ httpRequest.Content = new StringContent(
+ JsonSerializer.Serialize(request), Encoding.UTF8, "application/json");
+ using var response = await _http.SendAsync(
+ httpRequest, cancellationToken);
+ var body = await response.Content.ReadAsStringAsync();
+ if (!response.IsSuccessStatusCode)
+ {
+ var errorMessage = body;
+ try
+ {
+ using var errorDocument = JsonDocument.Parse(body);
+ if (errorDocument.RootElement.TryGetProperty("error", out var error)
+ && error.TryGetProperty("message", out var message))
+ {
+ errorMessage = message.GetString() ?? body;
+ }
+ }
+ catch
+ {
+ // 保留原始响应体。
+ }
+ return new TestAiResponse(
+ null, null, errorMessage, (int)response.StatusCode);
+ }
+
+ string? content = null;
+ string? reasoning = null;
+ using var document = JsonDocument.Parse(body);
+ AppendFromResponse(document.RootElement, ref content, ref reasoning);
+ return new TestAiResponse(
+ content, reasoning, null, (int)response.StatusCode);
+ }
+
+ private static void AppendFromResponse(
+ JsonElement root,
+ ref string? content,
+ ref string? reasoning)
+ {
+ if (!root.TryGetProperty("choices", out var choices)
+ || choices.ValueKind != JsonValueKind.Array
+ || choices.GetArrayLength() == 0)
+ {
+ return;
+ }
+ var message = choices[0].GetProperty("message");
+ if (message.TryGetProperty("content", out var contentProperty)
+ && contentProperty.ValueKind == JsonValueKind.String)
+ {
+ content = contentProperty.GetString();
+ }
+ if (message.TryGetProperty("reasoning_content", out var reasoningProperty)
+ && reasoningProperty.ValueKind == JsonValueKind.String)
+ {
+ reasoning = reasoningProperty.GetString();
+ }
+ }
+ }
+
+ internal static class OpenCodeGoFakeToolCallTests
+ {
+ private const string SystemPrompt =
+ "你是用于验证思维链回传的实验辅助。请先用 reasoning_content 进行内部推理,再给出简短、准确的回答。";
+ private const string FirstUserPrompt =
+ "请分析:A 比 B 高 20%,B 比 C 高 25%,那么 A 比 C 高多少?请先思考,再回答最终百分比。";
+ private const string ToolName = "analysis_hint";
+ private const string MemoryMarker = "TOKEN_MARK=731942";
+
+ public static void Run()
+ {
+ if (Environment.GetEnvironmentVariable("ARR_AI_E2E") != "1"
+ || Environment.GetEnvironmentVariable("ARR_AI_E2E_TOOL") != "1")
+ {
+ Console.WriteLine(
+ " [跳过] 未同时设置 ARR_AI_E2E=1 与 ARR_AI_E2E_TOOL=1,"
+ + "跳过伪造 tool call 历史实验");
+ return;
+ }
+ RunAsync().GetAwaiter().GetResult();
+ }
+
+ private static async Task RunAsync()
+ {
+ var settingsPath = FindSettingsPath();
+ Console.WriteLine($" [E2E-TOOL] 配置:{settingsPath}");
+
+ var settings = JsonSerializer.Deserialize(
+ File.ReadAllText(settingsPath, Encoding.UTF8))
+ ?? throw new InvalidOperationException("AI 设置文件无法解析");
+ var provider = settings.Providers.FirstOrDefault(p =>
+ p.Name.Equals("OpenCodeGo", StringComparison.OrdinalIgnoreCase)
+ && p.Models.Any(m =>
+ m.ModelId.Equals("deepseek-v4-flash", StringComparison.OrdinalIgnoreCase)))
+ ?? throw new InvalidOperationException(
+ "当前配置中未找到 OpenCodeGo / deepseek-v4-flash 服务与模型");
+ var model = provider.Models.First(m =>
+ m.ModelId.Equals("deepseek-v4-flash", StringComparison.OrdinalIgnoreCase));
+ if (string.IsNullOrWhiteSpace(provider.ApiKey))
+ {
+ throw new InvalidOperationException("OpenCodeGo 的 API Key 为空");
+ }
+
+ Console.WriteLine(
+ $" [E2E-TOOL] Provider={provider.Name}; Model={model.ModelId}; "
+ + $"BaseUrl={provider.BaseUrl}; Stream={model.IsStream}; ApiKey=***");
+
+ var report = new StringBuilder();
+ AppendHeader(report, provider, model, settingsPath);
+ var reportPath = Environment.GetEnvironmentVariable("ARR_AI_TOOL_REPORT_PATH");
+ if (string.IsNullOrWhiteSpace(reportPath))
+ {
+ reportPath = Path.Combine(
+ Environment.CurrentDirectory, "AI_reasoning_continuation_position_sweep_report.md");
+ }
+ reportPath = Path.GetFullPath(reportPath);
+
+ try
+ {
+ var analyzer = new AIAnalyze();
+ var normalContext = new AiRequestContext(provider, model);
+ var firstMessages = ImmutableList.Empty
+ .Add(new AIAnalyze.ChatMessage("system", SystemPrompt))
+ .Add(new AIAnalyze.ChatMessage("user", FirstUserPrompt));
+ var first = await CompleteAndRecordAsync(
+ analyzer, normalContext, firstMessages, "初始请求(system + user)", report);
+ var firstReasoning = first?.Reasoning ?? string.Empty;
+ if (first is null || firstReasoning.Length == 0)
+ {
+ report.AppendLine("初始请求未获得 reasoning_content,无法继续伪造 tool call 实验。");
+ return;
+ }
+ var initialFirst = first;
+
+ var toolContext = BuildToolContext(provider, model);
+ var wordings = new[]
+ {
+ new
+ {
+ Name = "措辞B:当前推理",
+ Target = "你当前正在进行的内部推理",
+ ArgumentInstruction = "继续你当前正在进行的内部推理,并参考工具结果。"
+ },
+ new
+ {
+ Name = "措辞D:tool_calls 携带的 reasoning_content",
+ Target = "你本次 tool_calls 消息中携带的 reasoning_content",
+ ArgumentInstruction = "继续你本次 tool_calls 消息中携带的 reasoning_content,并参考工具结果。"
+ }
+ };
+ var positions = new[]
+ {
+ new { Name = "25%", Ratio = 0.25 },
+ new { Name = "50%", Ratio = 0.50 },
+ new { Name = "75%", Ratio = 0.75 },
+ new { Name = "末尾", Ratio = 1.0 }
+ };
+ const int repetitions = 5;
+ var combinedCount = wordings.Length * positions.Length;
+ var foundCounts = new int[combinedCount];
+ var missingCounts = new int[combinedCount];
+ var failedCounts = new int[combinedCount];
+ var names = new string[combinedCount];
+ var markerInSent = new bool[combinedCount];
+ var comboIndex = 0;
+
+ foreach (var wording in wordings)
+ {
+ foreach (var position in positions)
+ {
+ var comboName = $"{wording.Name} / 标记位置 {position.Name}";
+ names[comboIndex] = comboName;
+ var reasoningWithMarker = InsertMemoryMarker(
+ firstReasoning, position.Ratio);
+ markerInSent[comboIndex] = reasoningWithMarker.Contains("731942");
+ Console.WriteLine(
+ $" [E2E-TOOL] 正在测试:{comboName}(重复 5 次)");
+ for (var rep = 0; rep < repetitions; ++rep)
+ {
+ var toolCallId = $"call_position_{comboIndex + 1}_{rep + 1}";
+ var arguments = JsonSerializer.Serialize(new
+ {
+ instruction = wording.ArgumentInstruction
+ });
+ var messages = ImmutableList.Empty
+ .Add(new AIAnalyze.ChatMessage("system", SystemPrompt))
+ .Add(new AIAnalyze.ChatMessage("user", FirstUserPrompt))
+ .Add(AIAnalyze.ChatMessage.AssistantToolCall(
+ reasoningWithMarker, toolCallId, ToolName, arguments))
+ .Add(AIAnalyze.ChatMessage.ToolResult(
+ toolCallId, BuildToolResult(wording.Target)));
+ var result = await CompleteAndRecordAsync(
+ analyzer, toolContext, messages,
+ $"记忆标记 {comboName}(重复 {rep + 1}/{repetitions})", report);
+ if (result is { } completed)
+ {
+ var combined = (completed.Reasoning ?? string.Empty)
+ + "\n" + completed.Response;
+ var found = ContainsMemoryMarker(combined);
+ var missing = SaysMarkerMissing(combined);
+ if (found)
+ {
+ foundCounts[comboIndex]++;
+ }
+ if (missing)
+ {
+ missingCounts[comboIndex]++;
+ }
+ Console.WriteLine(
+ $" [E2E-TOOL] {comboName} [{rep + 1}/{repetitions}]:"
+ + $"content={completed.Response.Length};"
+ + $"reasoning={(completed.Reasoning ?? string.Empty).Length};"
+ + $"markerFound={found};missing={missing}");
+ }
+ else
+ {
+ failedCounts[comboIndex]++;
+ Console.WriteLine(
+ $" [E2E-TOOL] {comboName} [{rep + 1}/{repetitions}]:请求失败(详见报告)");
+ }
+ }
+ comboIndex++;
+ }
+ }
+ AppendSummary(
+ report,
+ names,
+ markerInSent,
+ foundCounts,
+ missingCounts,
+ failedCounts);
+ }
+ finally
+ {
+ var reportDir = Path.GetDirectoryName(reportPath);
+ if (!string.IsNullOrEmpty(reportDir))
+ {
+ Directory.CreateDirectory(reportDir);
+ }
+ File.WriteAllText(reportPath, report.ToString(), new UTF8Encoding(false));
+ Console.WriteLine($" [E2E-TOOL] 人工审阅报告已写入:{reportPath}");
+ }
+ }
+
+ private static AiRequestContext BuildToolContext(
+ AiProvider provider,
+ AiModel sourceModel)
+ {
+ var model = new AiModel
+ {
+ ModelId = sourceModel.ModelId,
+ DisplayName = sourceModel.DisplayName,
+ IsStream = sourceModel.IsStream,
+ ContextLength = sourceModel.ContextLength,
+ ContextBudget = sourceModel.ContextBudget,
+ ExtraParameters = new Dictionary(sourceModel.ExtraParameters)
+ };
+ model.ExtraParameters["tools"] = new[]
+ {
+ new
+ {
+ type = "function",
+ function = new
+ {
+ name = ToolName,
+ description = "向模型提供上一轮思维链的继续提示。",
+ parameters = new
+ {
+ type = "object",
+ properties = new
+ {
+ instruction = new
+ {
+ type = "string",
+ description = "模型应遵循的续写提示。"
+ }
+ },
+ required = new[] { "instruction" }
+ }
+ }
+ }
+ };
+ model.ExtraParameters["tool_choice"] = "none";
+ return new AiRequestContext(provider, model);
+ }
+
+ private static string BuildToolResult(string target) =>
+ "工具结果:请继续" + target + "。这是一次确定性记忆测试。"
+ + "该思考内容中有一个内部记忆标记,请不要从本工具结果中猜测或寻找。"
+ + "请原样写出该标记;如果该思考内容中没有该标记,请明确回答“标记不存在”。"
+ + "不要从零重新推导,也不要杜撰标记。";
+
+ private static string InsertMemoryMarker(string text, double ratio)
+ {
+ if (ratio >= 1)
+ {
+ return text + "\n\n[内部记忆标记 " + MemoryMarker + "]";
+ }
+ var index = Math.Max(
+ 0,
+ Math.Min(text.Length, (int)(text.Length * ratio)));
+ return text.Insert(
+ index,
+ "\n\n[内部记忆标记 " + MemoryMarker + "]\n\n");
+ }
+
+ private static bool ContainsMemoryMarker(string text) =>
+ text.Contains("731942") || text.Contains("TOKEN_MARK");
+
+ private static bool SaysMarkerMissing(string text) =>
+ text.Contains("标记不存在")
+ || text.Contains("没有该标记")
+ || text.Contains("没有标记");
+
+ private static void AppendSummary(
+ StringBuilder report,
+ string[] names,
+ bool[] markerInSent,
+ int[] foundCounts,
+ int[] missingCounts,
+ int[] failedCounts)
+ {
+ report.AppendLine("## 记忆标记统计");
+ report.AppendLine();
+ report.AppendLine($"- 记忆标记:`{MemoryMarker}`");
+ report.AppendLine("- 每个组合重复 5 次。");
+ report.AppendLine();
+ for (var i = 0; i < names.Length; ++i)
+ {
+ report.AppendLine(
+ $"- **{names[i]}**:发送链包含标记 = {markerInSent[i]};"
+ + $"正确写出标记 = {foundCounts[i]}/5;"
+ + $"明确说没有 = {missingCounts[i]}/5;"
+ + $"请求失败 = {failedCounts[i]}/5");
+ }
+ report.AppendLine();
+ }
+
+ private static async Task CompleteAndRecordAsync(
+ AIAnalyze analyzer,
+ AiRequestContext context,
+ ImmutableList messages,
+ string label,
+ StringBuilder report)
+ {
+ var sw = Stopwatch.StartNew();
+ try
+ {
+ var result = await analyzer.CompleteAsync(
+ messages, context, null, CancellationToken.None);
+ sw.Stop();
+ AppendRecord(report, label, messages, result, sw.Elapsed);
+ return result;
+ }
+ catch (Exception ex)
+ {
+ sw.Stop();
+ report.AppendLine($"## {label}(请求失败)");
+ report.AppendLine();
+ report.AppendLine($"- 耗时:{sw.Elapsed.TotalSeconds:0.00} 秒");
+ report.AppendLine($"- 错误:{ex.Message}");
+ report.AppendLine("- 关键点:assistant 消息包含 tool_calls,随后有 tool 回复;请求携带 tools 定义。");
+ report.AppendLine();
+ report.AppendLine("---");
+ report.AppendLine();
+ Console.WriteLine($" [E2E-TOOL] {label} 失败:{ex.Message}");
+ return null;
+ }
+ }
+
+ private static void AppendHeader(
+ StringBuilder report,
+ AiProvider provider,
+ AiModel model,
+ string settingsPath)
+ {
+ report.AppendLine("# 伪造 tool call 历史实验");
+ report.AppendLine();
+ report.AppendLine($"- 配置:{provider.Name} / {model.ModelId}");
+ report.AppendLine($"- BaseUrl:{provider.BaseUrl}");
+ report.AppendLine($"- 配置文件:{settingsPath}");
+ report.AppendLine("- API Key:***(不写入报告)");
+ report.AppendLine();
+ report.AppendLine("本实验在请求中声明了 `tools` 并将 `tool_choice` 设为 `none`;");
+ report.AppendLine("构造的历史为:system → 初始 user → assistant(reasoning_content + tool_calls) → tool(tool_call_id + 提示结果)。");
+ report.AppendLine(
+ $"完整思维链按 25%/50%/75%/末尾四种位置注入确定性记忆标记 `{MemoryMarker}`;"
+ + "工具结果使用“当前推理”和“本次 tool_calls 消息携带的 reasoning_content”两种措辞,"
+ + "且不提及“上一轮”“截断点”等概念,要求模型原样写出标记。");
+ report.AppendLine("没有第 5 条新的 user 消息。");
+ report.AppendLine();
+ }
+
+ private static void AppendRecord(
+ StringBuilder report,
+ string label,
+ ImmutableList messages,
+ AIAnalyze.Result result,
+ TimeSpan elapsed)
+ {
+ report.AppendLine($"## {label}");
+ report.AppendLine();
+ report.AppendLine($"- 耗时:{elapsed.TotalSeconds:0.00} 秒");
+ report.AppendLine($"- prompt_tokens:{result.PromptTokens?.ToString() ?? "?"}");
+ report.AppendLine($"- completion_tokens:{result.CompletionTokens?.ToString() ?? "?"}");
+ report.AppendLine($"- reasoning_tokens:{result.ReasoningTokens?.ToString() ?? "?"}");
+ report.AppendLine();
+ report.AppendLine("### 发送的 messages");
+ report.AppendLine();
+ for (var i = 0; i < messages.Count; ++i)
+ {
+ var message = messages[i];
+ report.AppendLine($"**{i + 1}. role = `{message.Role}`**");
+ report.AppendLine();
+ if (message.Content is null)
+ {
+ report.AppendLine("content:`未发送`");
+ }
+ else
+ {
+ report.AppendLine("content:");
+ AppendTextBlock(report, message.Content);
+ }
+ report.AppendLine();
+ if (message.ToolCallId is not null)
+ {
+ report.AppendLine($"tool_call_id:`{message.ToolCallId}`");
+ report.AppendLine();
+ }
+ if (message.ToolCalls is { } toolCalls)
+ {
+ report.AppendLine("tool_calls:");
+ foreach (var toolCall in toolCalls)
+ {
+ report.AppendLine(
+ $"- id=`{toolCall.Id}`; type=`{toolCall.Type}`; "
+ + $"name=`{toolCall.Function.Name}`");
+ report.AppendLine();
+ report.AppendLine("arguments:");
+ AppendTextBlock(report, toolCall.Function.Arguments);
+ }
+ report.AppendLine();
+ }
+ if (message.Role == "assistant")
+ {
+ report.AppendLine(
+ message.ReasoningContent is null
+ ? "reasoning_content:`未发送`"
+ : "reasoning_content:");
+ if (message.ReasoningContent is not null)
+ {
+ AppendTextBlock(report, message.ReasoningContent);
+ }
+ report.AppendLine();
+ }
+ report.AppendLine();
+ }
+ report.AppendLine("### 模型响应");
+ report.AppendLine();
+ report.AppendLine("新 reasoning_content:");
+ AppendTextBlock(report, result.Reasoning ?? string.Empty);
+ report.AppendLine();
+ report.AppendLine("回答正文(content):");
+ AppendTextBlock(report, result.Response);
+ report.AppendLine();
+ report.AppendLine("---");
+ report.AppendLine();
+ }
+
+ private static void AppendTextBlock(StringBuilder report, string? text)
+ {
+ report.AppendLine("```text");
+ report.AppendLine(string.IsNullOrEmpty(text) ? "(空)" : text!);
+ report.AppendLine("```");
+ }
+
+ private static string FindSettingsPath()
+ {
+ var explicitPath = Environment.GetEnvironmentVariable("ARR_AI_SETTINGS_PATH");
+ if (!string.IsNullOrWhiteSpace(explicitPath) && File.Exists(explicitPath))
+ {
+ return explicitPath;
+ }
+
+ var dir = new DirectoryInfo(AppContext.BaseDirectory);
+ for (var i = 0; i < 6 && dir is not null; ++i)
+ {
+ var candidate = Path.Combine(
+ dir.FullName, "bin", "Debug", "net461", "AnotherReplayReader.ai_settings.json");
+ if (File.Exists(candidate))
+ {
+ return candidate;
+ }
+ var directCandidate = Path.Combine(
+ dir.FullName, "AnotherReplayReader.ai_settings.json");
+ if (File.Exists(directCandidate))
+ {
+ return directCandidate;
+ }
+ dir = dir.Parent;
+ }
+ throw new InvalidOperationException(
+ "未找到 AnotherReplayReader.ai_settings.json;可用 ARR_AI_SETTINGS_PATH 指定");
+ }
+ }
+
internal static class UserReplayFactIndexReproTests
{
public static void Run()
diff --git a/PLAN_ai_analysis_v2.md b/PLAN_ai_analysis_v2.md
index e0c05cc..ba3197f 100644
--- a/PLAN_ai_analysis_v2.md
+++ b/PLAN_ai_analysis_v2.md
@@ -2,13 +2,13 @@
## 状态
-- 日期:2026-08-20
-- 版本:v2(重写)。v1 初稿未覆盖第一轮审视的部分问题;本文新增第 2 节"问题追踪表"并重组章节结构,确保第一轮提出的每一条问题都有对应的落地章节。
-- 关联文档:[WIP.md](WIP.md)、[CONTEXT.md](CONTEXT.md)、`docs/adr/0001-hidden-revision-pass.md`、`docs/adr/0002-structured-game-knowledge.md`
+- 日期:2026-08-23
+- 版本:v2.1。WIP/CONTEXT/ADR 旧文档已删除;思维链回传实验单独记录在 `AI_REASONING_CONTINUATION_RESEARCH.md`。
+- 关联文档:[AI_REASONING_CONTINUATION_RESEARCH.md](AI_REASONING_CONTINUATION_RESEARCH.md)
## 实施状态(2026-08-20)
-里程碑全部完成,代码已落地并通过 124 项单元测试(`AiV2.Tests`,见 §12 M7)。
+里程碑全部完成,代码已落地并通过 149 项单元测试(`AiV2.Tests`,见 §12 M7;启用真实回放诊断时为 151 项)。
| 里程碑 | 状态 | 备注 |
| --- | --- | --- |
@@ -24,7 +24,7 @@
- `Data/StringHashes.xml` 是随仓库分发的本地 SDK 临时快照(约 3.5MB / 47,860 条),后续应改为可配置路径或只打包需要的 hash 子集。
- Corona 结构化知识(`knowledge_units_corona.json`)尚未编写:Corona 当前走 flat 文本(不剥离、不注入结构化条目),验证回退到启发式。
-- 修订 pass 的展示采用"实时流式 + 修订后整段替换";ADR 0001 的"完全缓冲至验证完成"仍是开放项。
+- 修订 pass 的展示采用"实时流式 + 修订后整段替换";原隐藏修订决策中的"完全缓冲至验证完成"仍是开放项。
- `Fatal` 在“所有机器可读声明块均无法解析”时产生;修订输出为空时保留原分析。
- `MissingMachineReadableClaims` 为 Warning,并与其他 Warning/WeakEvidence 一样触发一次隐藏修订;是否保留该策略待 A/B 评估。
- 测试工程 `AiV2.Tests` 通过 `ProjectReference` 引用主工程;构建时通过 `AiV2TestsBuilding=true` 跳过主工程的 DLL 移动目标。
@@ -44,10 +44,13 @@
- 机械分段超过上限时先过滤纯选择/编队事件块,压缩无效才放宽预算。
- Corona flat 文本现在也按参战阵营过滤。
- `AiV2.Tests` 的真实回放诊断改为默认跳过(设置 `ARR_E2E_REPLAY=1` 启用);UI 状态栏显示推理 token。
+- 思维链回传实验:主项目不再包含 `reasoning_content` 回传、截断、UI 实验开关或 tool call 历史构造代码,未进入生产管线。
+- `AiV2.Tests` 仅保留 `OpenCodeGoFakeToolCallE2e` 专用测试:默认跳过,需同时设置 `ARR_AI_E2E=1` 与 `ARR_AI_E2E_TOOL=1`。
+- 实验结果、推荐方案与数据表格见 `AI_REASONING_CONTINUATION_RESEARCH.md`。
## 1. 背景与目标
-应用现有 AI 分析流程为"全量日志 + LLM 分段建议 + 分段分析 + 总结",经文档与代码审视,存在四类问题:上下文膨胀(每轮重发全量日志)、验证层可信度(施法者/目标混淆、所有权证据缺失、解析脆弱)、知识层作用域(结构化数据无 mod 维度、提示词与验证知识漂移)、修订机制未落地(ADR 0001)。
+应用现有 AI 分析流程为"全量日志 + LLM 分段建议 + 分段分析 + 总结",经旧文档与代码审视,存在四类问题:上下文膨胀(每轮重发全量日志)、验证层可信度(施法者/目标混淆、所有权证据缺失、解析脆弱)、知识层作用域(结构化数据无 mod 维度、提示词与验证知识漂移)、修订机制未落地。
本计划的目标:
@@ -55,15 +58,15 @@
2. 只维护一条分析管线:"短录像 = 只有一个 slice",不保留两个独立模式。
3. 上下文预算成为每模型可配置的软上限,并把长期未使用的 `ContextLength` 接进护栏。
4. 修正验证层与知识层在本会话中发现的所有问题(见第 2 节追踪表)。
-5. 落地 ADR 0001 的隐藏修订 pass。
+5. 落地隐藏修订 pass。
## 2. 问题追踪表(第一轮审视 + 后续讨论确认)
-下表汇总 2026-08-20 会话对 WIP/文档/代码的审视结论。后续讨论(所有权分层、缓存、预算、统一管线)调整了部分原始结论,表中"处理章节"指向本文的落地位置。
+下表汇总 2026-08-20 会话对旧文档与代码的审视结论。后续讨论(所有权分层、缓存、预算、统一管线)调整了部分原始结论,表中"处理章节"指向本文的落地位置。
| # | 问题 | 处理章节 |
| --- | --- | --- |
-| P1 | 结构化单位知识无 mod 维度:`knowledge_units.json` 是全局单例,Corona 盟军数据被基础版替换(如 `AlliedBomberAircraft` vs `AlliedAntiStructureBomberAircraft`),违背 ADR 0002"每 mod 自包含" | §8.1 |
+| P1 | 结构化单位知识无 mod 维度:`knowledge_units.json` 是全局单例,Corona 盟军数据被基础版替换(如 `AlliedBomberAircraft` vs `AlliedAntiStructureBomberAircraft`),违背"每 mod 自包含"原则 | §8.1 |
| P2 | `GetSystemPrompt` 无条件执行旧 `BuildDefaultSystemPrompt`,即使走知识文件也会弹"缺乏苏联/未知地图"MessageBox(副作用) | §8.2 |
| P3 | 事实索引把特殊能力的施法者与目标混淆(`0x201/0x232` 的 ObjectId 不一定是施法者),Contradiction 校验可能误报/漏报 | §7.2 |
| P4 | 协议(`0x24E 选择协议`)不在验证体系,evidence schema 无法表达无单位能力 | §7.3 |
@@ -90,7 +93,7 @@
| 回查机制 | 进 v1;允许模型按需请求远处原始区间 |
| 缓存 | 稳定内容前置;跨段前缀 = system+摘要+总览;段内复用 = 前缀+slice(修订/回查共用) |
| 128K 及以下 | 允许短录像(切片后为 1 个 slice 时自然工作),不承诺长录像质量 |
-| 修订 pass | 按 ADR 0001 在段内落地,每段最多 1 次 |
+| 修订 pass | 按隐藏修订方案在段内落地,每段最多 1 次 |
## 4. 上下文预算策略
@@ -248,7 +251,7 @@
- flat 文本按参战阵营过滤非参战阵营章节;结构化渲染只渲染参战阵营(`RenderAsPrompt` 已有 factionNames 参数,flat 文本需要配套切分)。
-### 8.5 用户知识 JSON 加载(WIP 遗留)
+### 8.5 用户知识 JSON 加载(实施遗留)
- 落地 `AnotherReplayReader.user_knowledge.json`:按 id 覆盖内置条目,加载顺序:内置 → 用户覆盖。
@@ -268,12 +271,12 @@
- 补充 `protocol|时间|科技名` 类型。
- 注明 `move` 证据的验证限制(见 §7.5)。
-## 10. 修订 pass(ADR 0001 落地,P11)
+## 10. 修订 pass(P11)
- 段内执行:草稿 + 验证 issue + 相关事实 → 干净修正版;每段最多 1 次。
- 修订后仍 `Fatal` → 回退显示原文 + 警告(`Fatal` 条件见 §7.7)。
- 修订请求复用段内会话(同一前缀,缓存友好)。
-- UI:增加"验证器发现并修正 N 个问题"提示;沿用 ADR 0001 的缓冲决策(段内容缓冲至验证/修订完成)。
+- UI:增加"验证器发现并修正 N 个问题"提示;沿用隐藏修订方案的缓冲决策(段内容缓冲至验证/修订完成)。
## 11. 设置与 UI