# ADR 0002: Structured Game Knowledge for Prompt and Validation **Date:** 2026-07-07 ## Status Accepted ## Context Game knowledge — unit capabilities, faction rules, map geometry, build restrictions, and known exceptions — currently lives in two unconnected places: 1. **Prompt text** inside `AIAnalyze.BuildDefaultSystemPrompt()` as large hardcoded strings with `[MOD:]` conditional tags. 2. **Validation rules** inside `AIAnalysisValidation.cs` as hardcoded string matching (e.g., `claimLooksLikeBuilder` checks for `"MCV"`, `"基地车"`, `"Nanocore"`, etc.). This causes several problems: - Prompt knowledge and validation knowledge are not synchronized. Adding a new unit type requires editing both the prompt text and the validation code. - Users can only override the entire system prompt or append text. There is no way to add or correct a single unit fact without replacing the whole prompt. - Validation rules use fragile substring matching against natural-language Chinese text, which will drift as the prompt text changes. - There is no reusable data structure that both prompt rendering and validation logic can query. We need a unified knowledge architecture that: - Serves as the single source of truth for both prompt rendering and deterministic validation. - Lets users add map-, faction-, or mod-specific knowledge without editing code. - Replaces hardcoded substring matching in validation with tag-based queries. - Preserves the built-in game knowledge as the default for each supported mod. ## Decision ### Knowledge Set structure Game knowledge is organized into **KnowledgeSets**, each keyed by mod name (e.g., `"default"` for base game, `"corona"` for the Corona mod). Each set is **self-contained and complete** — there is no cross-set inheritance or conditional inclusion. The mod name from the replay directly selects which set to load, replacing the current `[MOD:]` inline tag system entirely. Each KnowledgeSet is a hierarchy where **scope is inherited from the path**, not stored in entries: ``` KnowledgeSet (e.g. "default") ├── global/ ← applies to all factions and maps │ ├── entries... ├── factions/ │ ├── 盟军/ │ │ ├── entries... ← scope = faction:盟军 │ ├── 神州/ │ │ ├── entries... ├── maps/ │ ├── map_mp_2_rao1/ │ ├── entries... ← scope = map:map_mp_2_rao1 ``` **Mod knowledge vs base game:** Since a mod is a self-contained game version, its KnowledgeSet is a complete copy of the relevant knowledge, not a diff. This avoids the complexity of conditional tags (`[MOD:]` / `[MOD:NO:]`) — users editing a mod's knowledge JSON see only that mod's entries without conditional logic. The current `[MOD:]` inline text approach is retired; knowledge sets are now purely data-driven. ### KnowledgeEntry format Every knowledge entry has the same structure whether it is built-in or user-supplied: ``` id: string # unique identifier within the knowledge set tags: string[] # from the predefined tag taxonomy (see below) text: string # markdown description, used for prompt rendering ``` ### Tag taxonomy (finite, predefined) Tags serve as the bridge between prompt knowledge and validation logic. Validation rules query entries by tag instead of matching strings. **Capability tags** (what a unit can do): `builder`, `pack`, `unpack`, `amphibious`, `transport`, `returnToProducer`, `cloak`, `toggleWeapon` **Type tags** (what a unit is): `infantry`, `vehicle`, `aircraft`, `naval`, `structure`, `hero`, `production`, `defense`, `superweapon` **Combat role tags** (what a unit fights): `antiInfantry`, `antiVehicle`, `antiStructure`, `antiAir`, `antiNaval` **Special power references** (links to observable replay data): `specialPower:PackReplaceSelf`, `specialPower:UnpackReplaceSelf`, etc. ### Prompt rendering The built-in `BuildDefaultSystemPrompt()` is refactored to render from the knowledge set in this order: 1. Global entries 2. Faction-specific entries for each player's faction (in player order) 3. Map-specific entries for the current map User customizations still apply as layers on top: - `AdditionalRules` is appended at the end of the rendered prompt. - `UseCustomSystemPrompt` completely replaces the default (as before). ### User extensibility Users can add or overlay knowledge entries via a JSON file (e.g., `AnotherReplayReader.user_knowledge.json`) stored alongside the settings file. The file follows the same KnowledgeSet structure; entries with matching `id` values override built-in entries. ### Validation consumption Validation rules (`AIAnalysisValidation.cs`) are refactored to: - Load the active knowledge set and query entries by tag (e.g., `entries.WithTag("builder")`) instead of matching hardcoded substrings. - Use `specialPower:*` tags to verify power-to-unit-type inferences. - Keep deterministic rules (e.g., `ValidateUnpackAmbiguity`) as code, but drive what entries they check from tags rather than hardcoded asset names. ### Storage format JSON container with text fields as markdown. Example: ```json { "global": { "entries": [ { "id": "general-rules", "tags": ["rule"], "text": "## 核心原则\n- ..." } ] }, "factions": { "盟军": { "entries": [ { "id": "AlliedMCV", "tags": ["builder", "vehicle", "amphibious", "pack", "unpack"], "text": "### 基地车(AlliedMCV)\n盟军基地车,两栖,...\n可以在陆地或水上展开为主基地。" } ] } } } ``` ## Consequences - Positive: Prompt knowledge and validation knowledge share a single source of truth. - Positive: Users can add or correct game knowledge facts without modifying code. - Positive: Validation no longer depends on fragile Chinese substring matching — tag queries are deterministic and language-independent. - Positive: New mods (e.g., "corona") get their own KnowledgeSet without polluting the default. - Negative: Requires migration of ~700 lines of hardcoded prompt text into KnowledgeEntry records — a significant one-time refactoring effort. - Negative: The tag taxonomy must be maintained as the game or mod evolves. Additions must be reviewed to prevent tag proliferation. - Negative: JSON file editing is less user-friendly than a dedicated settings UI (acceptable as an initial step; the `AiPromptSettings` text fields remain available as a simpler escape hatch). ## Implementation Status Not yet implemented. The following migration path is planned: 1. Define C# records (`KnowledgeSet`, `KnowledgeEntry`, `KnowledgeTag` constants) in a new `Utils/AiKnowledge.cs` file. 2. Create the built-in `KnowledgeSet` by extracting data from the current `BuildDefaultSystemPrompt()` strings into structured entries with tags. 3. Wire the knowledge set through `GetSystemPrompt()` so it renders entries in the correct order. 4. Refactor `AIAnalysisValidation.ValidateTimelineConsistency()` to query the knowledge set by tag instead of hardcoded matching. 5. Add user knowledge file loading in `AiSettings.Load()`.