# OFFICE OF METHOD ## O·M–04 / WP–001 # THE MIRAGE OF MASTERY ## An Evidence Architecture for Chess Skill Under Cue Withdrawal, Transfer, and Serious Competition **J. Bentley Tar** **Office of Method — Independent Laboratory** **CAPTUREDMIRAGE / Learning** **September 2026** **Version:** 1.0 — Pre-Field Working Paper **Research status:** Methodological paper, instrument specification, and prospective field-evaluation protocol; no prospective field results are reported in this version. **Implementation authority:** CAPTUREDMIRAGE 10.1.0 — TOURNAMENT FREEZE, semantic-content epoch 10.0. **Prospective field study:** Field Event A — a six-round, in-person classical Swiss tournament under a classical time control with increment. Exact event identity, schedule, time control, venue, section, and participant-administrative identifiers are retained in an embargoed, externally timestamped Pre-Event Lock rather than disclosed in the public manuscript. **Epistemic status:** CAPTUREDMIRAGE 10.1.0 is a frozen software research system whose specified internal release-assurance gates passed. Its Player Model, mastery stages, process diagnoses, transfer-warning heuristic, repair rules, and training policies have not been established as psychometrically validated measures or as interventions with demonstrated causal efficacy. **Canonical URL:** `https://officeofmethod.com/publications/the-mirage-of-mastery/` **Rights:** © 2026 J. Bentley Tar. All rights reserved. **Release identity:** `OM-04-WP-001-v1.0-MANIFEST.json` records the SHA-256 identities of the canonical TXT, PDF, and HTML publication artifacts. The manuscript does not embed a self-hash. **Version policy:** Version 1.0 remains archived unchanged after public release. Any substantive pre-field correction is issued as a separately versioned correction with the previous version retained. Field results will not silently replace this pre-outcome record. ### Suggested citation Tar, J. B. (2026). *The Mirage of Mastery: An Evidence Architecture for Chess Skill Under Cue Withdrawal, Transfer, and Serious Competition*. Office of Method Working Paper O·M–04/WP–001, Version 1.0. The term **mastery** is used throughout this paper as the object of an evidentiary claim. It does not denote a unitary latent psychological state assumed to exist independently of task, support, occasion, context, or measurement procedure. --- ## Abstract Chess training creates a measurement problem because the instructional environment can make successful performance easier to produce. A tactical exercise announces that a tactic is likely to exist. A strategic category identifies a family of considerations before the player has had to recognize it independently. Repetition converts an unfamiliar decision into a partly mnemonic task. Immediate correction can make subsequent success evidence of recent instruction rather than independent availability. These observations remain useful, but their evidentiary scope is narrower than a claim about performance after delay, discrimination, cue withdrawal, source variation, modality change, opponent resistance, or serious competition. CAPTUREDMIRAGE is a longitudinal chess-development system built to preserve these distinctions before operational scoring collapses them. Its frozen 10.1.0 release records dimension-specific bounded evidence, provenance, instructional context, prompting, discrimination status, source identity, recency, physical-board transfer, serious-game provenance, and—in designated modules—decision-process information. It separates supported semantics from quarantined semantics, active material from sealed assessment material, instruction from independent retest, and software-release assurance from human-performance efficacy. Its calculation system decomposes human search into orienting, candidate generation, resistance modeling, branch conclusion, candidate comparison, checking, and commitment. Its repair system does not allow immediate instructional success to certify an independently tested repair. This paper distinguishes two analytic planes. The **Operational Plane** uses heuristic context weights, quality coefficients, recency discounting, pseudo-counts, and practical thresholds to allocate training online. The **Research Plane** evaluates claims and therefore does not inherit those conveniences automatically. It requires native outcome identity, explicit dimension attribution, common scoring families within estimands, decomposed evidence facets, declared target performance domains, outcome-independent sampling, explicit opportunity denominators, process-provenance restrictions, and versioned analysis rules. The distinction is mathematically consequential. Under the frozen accumulator form, the operational Transfer Gap can be non-zero even when the observed bounded performance score is identical in supported and transfer contexts, solely because different evidence masses shrink by different amounts toward the same symmetric pseudo-prior. The paper derives this unequal-shrinkage result explicitly. The shipped Transfer Gap can therefore remain useful as an operational warning while being inadmissible as a scientific estimator of transfer loss. The first major prospective ecological evaluation is Field Event A. It is treated as a single-participant field evaluation and protocol stress test, not an efficacy trial. Pre-event warnings are connected to tournament evidence only through a frozen Field Criterion Map. The primary prospective question is framed as a **prospective warning–criterion profile**, not psychometric validation: for each dimension that can be assessed under the field protocol, what pre-event warning state and later serious-game success/failure/indeterminate profile are observed under the frozen field criteria? No common cross-dimensional concordance coefficient is assumed. Candidate omission cannot be inferred merely from retrospective silence; engine adjudication evaluates every explicitly remembered candidate even when it falls outside the initial MultiPV output; and missing process evidence remains indeterminate rather than being silently scored as success or failure. The governing restriction is: > **Evidence obtained under one set of supports should not settle a competence claim whose scope depends on performance after those supports have been removed or changed.** --- # 1. INTRODUCTION Chess appears unusually measurable because much of its external state can be recovered. Games can be recorded move by move, positions can be reconstructed exactly, legal alternatives can be regenerated, and fixed-budget engine analysis can provide a reproducible computational reference for concrete move quality. Eligible endgames can be evaluated with tablebases under stated conventions. Competitive performance can be summarized by established rating systems. Compared with many complex skills, chess leaves an unusually rich external trace. The trace does not identify the competence that produced it. A strong move may result from accurate calculation, recognition of a known pattern, memory of a previous position, successful strategic abstraction, or an unsound process that happened to terminate at a good answer. A poor move may arise because the player never generated a relevant candidate, represented the board incorrectly, modeled weaker resistance than the opponent could plausibly provide, terminated a branch too early, evaluated a correctly reached endpoint incorrectly, confused move order, misjudged the strategic transformation, or made a coherent practical decision that happened to encounter one concrete refutation. The move is externally observable. The pathway that produced it is only partially observable. Training adds another ambiguity because the learning environment modifies the problem before the player encounters it. A position placed under *back-rank tactics* no longer tests whether the player independently recognizes that a back-rank mechanism deserves attention. A position presented under *minority attack* reduces the discrimination problem before strategic reasoning begins. A repertoire trainer that has displayed the same sequence repeatedly can measure retrieval accurately while saying much less about orientation after an unfamiliar deviation. These are legitimate instructional conditions. They become problematic only when success inside them is interpreted as though those supports were absent. Transfer research has long resisted such simplification. Barnett and Ceci (2002) treat transfer as varying across several dimensions rather than along one universal near-to-far scale. Morris, Bransford, and Franks (1977) show that retrieval depends partly on correspondence between the operations engaged during learning and those demanded later. A decrement after context change can therefore indicate incomplete generalization, but it can also arise because the later condition genuinely imposes a different task. The correct question is not whether performance is numerically invariant under every transformation. It is whether the evidence adequately represents the conditions over which the competence claim is intended to range. Generalizability Theory supplies a complementary constraint. Observed performance is sampled under facets such as tasks, occasions, methods, and raters; task sampling itself can account for substantial variability in performance assessment (Shavelson & Webb, 1991; Shavelson et al., 1993). Two chess positions tagged `candidate-generation` can differ dramatically in difficulty and in the other competencies they engage. A simple forcing position and a rich middlegame containing several strategically plausible moves do not become interchangeable measurement occasions because a taxonomy gives them the same label. Chess expertise also changes the form of the problem. Classic work on chess memory and perception demonstrates that strong players encode meaningful configurations structurally rather than as unrelated piece locations (Chase & Simon, 1973; de Groot, 1978; Gobet & Simon, 1996a, 1996b, 1996c). Yet stored structure can itself mislead. Bilalić, McLeod, and Gobet (2008a, 2008b) demonstrated that familiar good solutions can inhibit discovery of better ones: expertise can produce an Einstellung effect rather than merely protect against error. A development system therefore cannot equate more pattern recognition with better cognition. It must also care about discrimination, alternative generation, resistance, and the conditions under which a familiar representation should be abandoned. CAPTUREDMIRAGE developed as an operational system before every distinction in this paper had been formalized. The frozen release nevertheless preserves substantially more information than the Player Model ultimately compresses. Its evidence events retain a modeled dimension, bounded score where one exists, context, quality, prompting, discrimination status, reliability, source, source identity, metadata, and time. Physical-board sessions preserve whether the digital position was revealed. Serious-game Clinic records can generate dimension-specific evidence and repair cases. Sealed benchmark state is retained separately from ordinary training. That asymmetry is valuable. It means an operational scoring rule can be criticized or replaced without having destroyed every distinction needed to criticize it. The paper therefore separates two analytic planes. The **Operational Plane** is an online control system. It must choose work under uncertainty and limited time. It may therefore use heuristic scalarization, recency discounting, pseudo-priors, thresholds, and approximate uncertainty if these mechanisms produce useful, bounded training allocation. The **Research Plane** has a different objective. It asks what a body of evidence supports. It therefore requires explicit estimands, scoring discipline, sampling design, dimension mappings, opportunity denominators, dependence accounting, contamination control, and prospective analysis rules. A quantity can belong legitimately to one plane without belonging to the other. This becomes particularly important because the paper identifies a structural defect in the interpretation suggested by the frozen operational Transfer Gap. The defect does not imply that the software is incorrectly implemented. It shows that a heuristic can perform an operational warning function while failing as an estimator of the scientific construct that its convenient name appears to denote. The purpose of this paper is consequently not to prove that CAPTUREDMIRAGE improves chess strength. The current software-assurance record cannot establish that proposition, and one six-game field event cannot establish it either. The purpose is to specify the instrument, articulate a theory of evidence adequate to its intended use, expose where the current operational mathematics cannot support stronger interpretation, and define a prospective procedure under which the system can begin encountering serious-game evidence that its training environment did not create. --- # 2. FROM TASK OUTCOMES TO COMPETENCE CLAIMS ## 2.1 Dimension-specific evidence events CAPTUREDMIRAGE 10.1.0 stores evidence at the level of a modeled dimension. The implementation creates an evidence event containing a `dimension`, a bounded `score` when one is supplied, an optional success flag, quality and context labels, prompting and discrimination flags, reliability, source provenance, metadata, and timestamp. The Player Model groups events by `dimension` before applying the operational accumulator (Office of Method, 2026d). The paper represents such an event as $$ e_i = \left( j_i, d_i, R_i, P_i, S_i, t_i, \mathbf{x}_i, \mathbf{m}_i \right), \tag{1} $$ where \(j_i\) identifies the underlying item or decision, \(d_i\) is the modeled dimension credited or debited by the event, \(R_i\) is the native task outcome, \(P_i\) is process evidence when such evidence exists, \(S_i\in[0,1]\) is the bounded operational evidence score used by the Player Model when defined, \(t_i\) is observation time, \(\mathbf{x}_i\) contains evidentiary facets, and \(\mathbf{m}_i\) contains measurement provenance. For tasks without process evidence, $$ P_i=\varnothing. \tag{2} $$ The distinction among \(R_i\), \(P_i\), and \(S_i\) is necessary. A binary tactical outcome, an ordinal branch-quality judgment, a reconstruction error, and a serious-game Clinic diagnosis are not the same variable. CAPTUREDMIRAGE may convert different observations into bounded operational evidence because the scheduler requires a common control interface. That conversion is an operational coding convention. Mapping an ordinal rubric into \([0,1]\) does not, by itself, establish interval-scale measurement. Nor does one task outcome automatically become evidence for every skill implicated by the position. One wrong move can coexist with correct board representation, adequate candidate generation, correct identification of the opponent's strongest resource, and a later endpoint-evaluation error. Conversely, one correct move can coexist with a defective process. An underlying position can therefore support several dimension-specific evidence events only when each event has a separately defensible basis. The system must not create pseudo-replication by copying one final result across every dimension that happens to be associated with the item. ## 2.2 Item-to-skill relevance and score admissibility Let $$ Q_{jd}\in\{0,1\} \tag{3} $$ denote whether item \(j\) is mapped, under a specified item-to-skill codebook, to dimension \(d\) as substantively relevant. This is a research mapping, not a claim that all mapped dimensions were actually observed in a particular encounter. For a dimension-specific score to be admissible, $$ S_i \text{ may contribute to dimension }d $$ only when $$ d_i=d, \qquad Q_{j_i d}=1, \tag{4} $$ and the event record contains sufficient information to apply the governing scoring rubric. Relevance therefore does not imply observability. A tournament position may clearly make candidate generation relevant while the post-game process record is too fragmentary to establish which candidates were considered. In that case the Q-matrix can mark candidate generation as relevant while no primary process score for candidate omission is admissible. This distinction is central to the field protocol. ## 2.3 Formalism-to-implementation correspondence The mathematical notation is intended to expose the logic of the release, not to imply that every symbol is the literal name of a storage field. ### Table 1. Formal quantities and CAPTUREDMIRAGE 10.1.0 implementation correspondence | Paper quantity | Implementation correspondence | Status | | -------------- | --------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------- | | \(j_i\) | Source/item/decision identity, typically represented through `sourceId`, task state, FEN, Clinic decision identity, or module-specific record | Implemented through module-specific identifiers | | \(d_i\) | `dimension` | Literal v10.1 evidence field | | \(R_i\) | Native task result before research interpretation | Module-specific; may be represented in task history, process record, metadata, or success/result fields | | \(P_i\) | Candidate lists, branch records, conclusions, process diagnoses, Clinic reconstruction, or analogous process content | Implemented where the module records process; not universal | | \(S_i\) | `score`, with success converted to 1/0 when score is absent in the core accumulator | Literal bounded operational input to the Player Model | | context | `context` | Literal evidence field | | quality | `quality` | Literal evidence field | | prompted | `prompted` | Literal evidence field | | discrimination | `discrimination` | Literal evidence field | | reliability | `reliability` | Literal evidence field; clamped by the core algorithm | | \(\omega_i\) | Optional event `weight` consumed by the core algorithm; first-party v10 event creation ordinarily leaves it at the implicit default | Implementation-supported optional modifier | | \(t_i\) | `at` | Literal timestamp field | The mathematical representation is therefore neither a fictitious replacement schema nor a claim that all module data are homogeneous. It is an analytical representation of the objects actually consumed by the frozen Player Model. ## 2.4 Target performance domains A competence claim requires a declared universe of admissible observations. For modeled dimension \(d\), define a target performance domain $$ \mathcal{D}_d = (\mathcal{G}_d,\Pi_d), \tag{5} $$ where \(\mathcal{G}_d\) is the set of task and context conditions over which the claim is intended to generalize and \(\Pi_d\) is the target distribution or weighting over those conditions. The sampling design used to collect evidence is represented separately as $$ \Sigma_d. \tag{6} $$ The distinction is substantive. \(\Pi_d\) describes the performance universe that matters; \(\Sigma_d\) describes how observations were actually selected. A study may oversample a rare but important condition. That does not alter the target population; it alters the sample. A claim specification is written $$ C_d = (\mathcal{D}_d,\tau_d,\mathcal{E}_d), \tag{7} $$ where \(\tau_d\) is the criterion relevant to the particular claim and \(\mathcal{E}_d\) is the minimum evidentiary requirement for treating the claim as supported rather than indeterminate. This formulation gives precise meaning to **reproducible availability**. Reproducibility does not mean that a skill appeared twice. It means that a declared evidence requirement has been met under conditions capable of representing the claim's target domain. ## 2.5 Worked example: candidate generation Consider the modeled dimension $$ d=\mathrm{CG}, $$ candidate generation. A broad long-run target domain might include unfamiliar, cue-free chess decisions in which more than one plausible move requires consideration and where the player must generate candidate moves without being told the tactical or strategic theme in advance. That long-run domain is too broad to be established by one tournament. For Field Event A, define the event-bounded target $$ \mathcal{G}^{A}_{\mathrm{CG}} = \{ \text{sampled Field Event A decisions mapped to candidate generation under the frozen task-demand codebook} \}. \tag{8} $$ The event-specific sampling design is $$ \Sigma^{A}_{\mathrm{CG}} = \text{the outcome-independent Corpus B procedure in Section 9.5}. \tag{9} $$ The field criterion is not simply whether the move played was engine top-1. A candidate-generation failure requires a process record sufficiently complete to support a negative inference and a candidate-complete reference analysis showing that none of the explicitly reconstructed candidate moves fell within the frozen acceptable-candidate set. The first field event can therefore support a narrow statement such as: > within the sampled event environment, candidate-generation opportunities produced a specified number of scorable successes, failures, and indeterminate cases under the frozen codebook. It cannot establish: > candidate generation is mastered across serious chess. The formal distinction between \(\mathcal{G}_{\mathrm{CG}}\) and \(\mathcal{G}^{A}_{\mathrm{CG}}\) prevents one bounded tournament from being mistaken for the entire performance domain. ## 2.6 Evidence facets rather than one natural context ladder The frozen v10.1 implementation uses one compound operational context label with the values `guided`, `curriculum`, `mixed`, `otb`, `sparring`, and `serious-game`. Those values do not constitute one natural scientific dimension. `guided`, `curriculum`, and `mixed` principally encode instructional support; `otb` principally identifies presentation modality; `sparring` identifies an opponent or practice environment; and `serious-game` identifies competitive consequence. A serious tournament game can simultaneously be physical, opponent-driven, cue-free, first-exposure, time-limited, rated, and competitively consequential. A physical training game can be both OTB and sparring. The Research Plane therefore represents an evidence-facet profile schematically as $$ \mathbf{x}_i = ( x_{\mathrm{support}}, x_{\mathrm{cue}}, x_{\mathrm{exposure}}, x_{\mathrm{modality}}, x_{\mathrm{opponent}}, x_{\mathrm{competition}}, x_{\mathrm{occasion}} ). \tag{10} $$ These are not assumed to form a metric space, and several are nominal. A guided task can isolate one mechanism cleanly while possessing weak ecological resemblance to competition. A sealed benchmark can have strong contamination control without opponent agency. A serious game possesses high ecological consequence while being comparatively poor at isolating one latent mechanism. The purpose of the Research Plane is therefore not to identify a universally superior evidence category. It is to preserve the conditions required to interpret the claim being made. ## 2.7 Operational mastery stages CAPTUREDMIRAGE's six frozen mastery stages are $$ \text{acquisition} \rightarrow \text{retention} \rightarrow \text{discrimination} \rightarrow \text{mixed transfer} \rightarrow \text{cue-free transfer} \rightarrow \text{serious-game evidence}. \tag{11} $$ These remain useful as operational graduation conditions. They should not be interpreted as six universally ordered psychological states. A trivial recapture in a tournament can be less diagnostically demanding than a difficult sealed discrimination task. A cue-free position may test independent recognition more strongly than an explicitly themed exercise while still lacking the clock, opponent agency, and stakes of competition. The appropriate inference rule is therefore applied to evidence sets: > **An evidence set supports a competence claim only to the extent that its task distribution, sampling design, support conditions, exposure state, modality, occasions, and measurement procedures identify the target performance domain specified for that claim.** --- # 3. THE FROZEN CAPTUREDMIRAGE 10.1.0 SYSTEM CAPTUREDMIRAGE 10.1.0 — TOURNAMENT FREEZE was frozen on 2 September 2026 over semantic-content epoch 10.0. Its exact executable identity, curriculum manifest, core algorithm identity, and release-assurance records are retained as Office of Method primary research artifacts (Office of Method, 2026a, 2026b, 2026c, 2026d). ### Table 2. Frozen CAPTUREDMIRAGE 10.1.0 curriculum scope | Component | Frozen scope | | --------------------------------------- | -------------------------------: | | Structural Intelligence Graph | 28 structures across 7 families | | Positional Attention Layer | 493 positions / 815 interactions | | Strategy Atlas | 1,200 positions | | Strategy Atlas — active | 493 | | Strategy Atlas — unsupported quarantine | 467 | | Strategy Atlas — sealed reserve | 240 | | Calculation Forge 5 | 240 positions | | Calculation failure taxonomy | 24 operational classes | | Endgame mechanisms | 32 | | Exact Endgame Oracle | 1,200 positions | | Exact-active WDL/DTZ cases | 171 | | Endgame Transition School | 72 positions | | Opening repertoires | 4 | | Opening lines | 48 | | Opening decisions | 635 | | Opening DAG nodes | 1,016 | | Historical Master Canon | 360 games | | Historical decision events | 1,361 | | Benchmark Core | 3 epochs / 42 forms / 300 items | In this table, an **interaction** is a stored position-specific relational record in the Positional Attention Layer; a **historical decision event** is a mapped decision point extracted from the historical-game corpus; an **exact-active case** is a currently active position with exact WDL/DTZ evidence under the configured tablebase boundary; and a **benchmark form** is a reserved assessment form belonging to one Benchmark Core epoch. The scale of the corpus is secondary to the distinctions encoded within it. A strategic position can exist in the corpus without being authorized for scored strategic instruction. Historical evidence can remain historical rather than being converted into an asserted causal explanation. A sealed item can retain first-exposure assessment value because it has not been permitted to leak into ordinary training. Exact tablebase evidence can correct an objective outcome claim without becoming an explanation of why the human player failed. ## 3.1 Provenance and semantic restraint The architecture distinguishes board-derived facts, human-curated strategic priors, historical evidence, engine verification, exact tablebase information where applicable, and human-model evidence. Stockfish supplies concrete move evidence under specified computational conditions; it is not treated as the author of human strategic meaning. Historical master moves establish what was played; they do not prove that one retrospective explanation uniquely caused the decision. Tablebases can establish convention-bound WDL/DTZ properties without identifying the most useful human teaching model. The strongest architectural rule is that unsupported strategic semantics may remain unsupported. No complete explanatory layer is manufactured merely because an empty field would be aesthetically inconvenient. Unknown is an admissible information state. ## 3.2 Semantic state and exposure state The Research Plane distinguishes two conceptual state axes. Semantic status is represented as $$ S_j^{\mathrm{sem}} \in \{ \text{supported}, \text{quarantined}, \text{historical-only} \}, \tag{12} $$ while exposure status is represented as $$ S_j^{\mathrm{exp}} \in \{ \text{active}, \text{sealed}, \text{consumed} \}. \tag{13} $$ The distinction matters because semantic support and exposure are orthogonal. A position can be semantically supported while remaining sealed for assessment. Exposure can consume first-exposure value without invalidating the position's underlying semantic record. Historical-only material can be used descriptively without becoming authorized scored strategic evidence. Where the frozen application stores some combinations through legacy state structures rather than two literal independent fields, Equations (12) and (13) represent the Research Plane's conceptual decomposition rather than a claim about the storage schema. ## 3.3 Release assurance All 28 enumerated current-state executable QA programs or contracts in the frozen release passed. A separate managed-browser navigation check remained outside the executable gate under environment policy. The 180-day synthetic run completed with 18,000 evidence events across 43 modeled dimensions and seven domains. Regression testing exercised physical-board authority, digital-reveal downgrade, setup-time exclusion, repair-stage gating, sealed-form consumption, disjoint active/sealed pools, and durable evidence continuity (Office of Method, 2026b). These results support software-system claims. They do not establish that the Player Model corresponds to a validated latent construct, that its internal thresholds are calibrated to human chess skill, or that using CAPTUREDMIRAGE improves competitive performance. Two separate builds under the frozen build procedure produced identical shipping executable hashes. This establishes deterministic build repeatability under the recorded procedure. It is not independent external reproduction. --- # 4. THE OPERATIONAL PLANE The Operational Plane exists to act under uncertainty. Its practical objective is to select useful work, prioritize repairs, prevent stale or heavily prompted evidence from dominating the present state, and alter training allocation as competition approaches. ### Table 3. Operational and Research analytic planes | Property | Operational Plane | Research Plane | | ---------------- | ------------------------------------- | ------------------------------------------------------------------------- | | Purpose | Online training allocation | Claim evaluation | | Time regime | Continuous | Frozen/versioned extract | | Outcome handling | Bounded operational scores permitted | Native scale preserved; common scoring family required within an estimand | | Context | Compound scalarization permitted | Facets decomposed | | Context weights | Heuristic policy | Excluded from primary research estimands unless independently justified | | Recency | Policy discounting permitted | Explicitly modeled, stratified, or reported | | Pseudo-prior | Stabilization permitted | Used only under an explicit probabilistic model | | Item composition | Heterogeneity tolerated operationally | Controlled, stratified, modeled, or declared | | Overrides | Logged operational action | No silent post hoc reweighting | | Main question | What should be trained next? | What does the evidence support? | ## 4.1 Frozen operational weighting The v10.1 context coefficients are $$ q_c= \begin{cases} 0.42 & \text{guided},\\ 0.62 & \text{curriculum},\\ 0.76 & \text{mixed},\\ 0.88 & \text{OTB},\\ 0.92 & \text{sparring},\\ 1.00 & \text{serious game}. \end{cases} \tag{14} $$ Quality grades A through E receive $$ q_g \in \{1.00,\ 0.84,\ 0.66,\ 0.46,\ 0.25\}. \tag{15} $$ Prompted evidence receives a multiplier of \(0.78\) relative to otherwise equivalent unprompted evidence. Reliability is clamped by the frozen core to the interval $$ r_i\in[0.1,1]. \tag{16} $$ The default recency half-life is 90 days, with a minimum configured half-life of 21 days. The optional event `weight` enters with a floor of \(0.25\), and the final event weight receives an overall floor of \(0.03\) in the core algorithm (Office of Method, 2026d). A schematic form of the implemented weight is therefore $$ w_i = \max \left[ 0.03,\; q_{c_i} q_{g_i} r_i 2^{-a_i/h_i} p_i \max(0.25,\omega_i) \right], \tag{17} $$ where \(a_i\) is evidence age, \(h_i\ge21\) is the configured half-life, \(p_i=0.78\) for prompted evidence and \(1\) otherwise, and \(\omega_i\) denotes the optional implementation-supported event-weight term. First-party v10 evidence creation ordinarily leaves that optional term at its implicit default. Equation (17) is an engineering composition rule. It assumes multiplicative separability. Prompting can interact with context, evidence quality can interact with reliability, and recency can matter differently for opening recall and strategic discrimination. Neither multiplicativity nor absence of interaction is established as a validated psychological model. ## 4.2 Dimension-specific weighted accumulation Let $$ \mathcal{I}_d = \{ i: d_i=d \text{ and } S_i \text{ is defined} \}. \tag{18} $$ The frozen beta-style accumulator can then be represented exactly at the level relevant to this paper as $$ \alpha_d = 2+ \sum_{i\in\mathcal{I}_d} w_iS_i, \qquad \beta_d = 2+ \sum_{i\in\mathcal{I}_d} w_i(1-S_i), \tag{19} $$ with operational evidence score $$ \mu_d = \frac{\alpha_d} {\alpha_d+\beta_d}. \tag{20} $$ Where a numeric `score` is absent, the core uses `success` as a binary fallback. The construction is a weighted operational evidence accumulator rather than a calibrated Bayesian posterior. Observations need not be independent or exchangeable, item difficulty is heterogeneous, repeated tasks can be related, the weight function is heuristic, and evidence contribution changes with time. The interval and uncertainty quantities returned by the same beta form are therefore operational uncertainty indicators, not validated frequentist confidence intervals or posterior credible intervals for latent chess ability. ## 4.3 Prior sensitivity For a set of dimension-specific events with total operational evidence mass $$ W = \sum_i w_i \tag{21} $$ and weighted observed bounded score $$ p = \frac{\sum_iw_iS_i}{W}, \qquad W>0, \tag{22} $$ a symmetric pseudo-prior \((a,a)\) yields $$ \mu_a(W;p) = \frac{a+pW}{2a+W}. \tag{23} $$ Table 4 illustrates the effect of prior mass. ### Table 4. Analytical prior-mass sensitivity | Observed weighted score \(p\) | Evidence mass \(W\) | \(a=1\) | \(a=2\) | \(a=5\) | | ----------------------------: | ------------------: | ------: | ------: | ------: | | 0.25 | 1 | 0.417 | 0.450 | 0.477 | | 0.25 | 3 | 0.350 | 0.393 | 0.442 | | 0.25 | 10 | 0.292 | 0.321 | 0.375 | | 0.25 | 30 | 0.266 | 0.279 | 0.313 | | 0.75 | 1 | 0.583 | 0.550 | 0.523 | | 0.75 | 3 | 0.650 | 0.607 | 0.558 | | 0.75 | 10 | 0.708 | 0.679 | 0.625 | | 0.75 | 30 | 0.734 | 0.721 | 0.688 | **Note.** These values are analytical illustrations, not empirical observations from the participant. Prior choice can dominate sparse operational evidence and remain materially influential at moderate evidence mass. The frozen \((2,2)\) prior is therefore an active control parameter, not an innocuous mathematical decoration. The Research Plane does not inherit it by default. ## 4.4 Recency is evidence reversion, not a forgetting law If no new evidence enters and all historical event weights decay toward zero, $$ \lim_{t\rightarrow\infty} \mu_d(t) = 0.5. \tag{24} $$ A high operational score therefore falls toward 0.5 while a low score rises toward 0.5. The mechanism does not model psychological forgetting directly. Three objects must remain distinct: the immutable historical evidence ledger; the present recency-weighted operational state; and any future validated learning/forgetting model. The 90-day half-life belongs to the second object. ## 4.5 Daily Conductor Daily Conductor 6 uses the Player Model, open repairs, due retrieval, uncertainty, prior external work, available time, and tournament phase to generate a bounded mission. The present paper does not reproduce an incomplete scheduler equation whose uncertainty term is insufficiently specified in the release documentation. The scientifically defensible claim is narrower: the scheduler treats both low estimated performance and insufficient knowledge of a dimension as reasons to allocate attention. Training therefore has two operational functions. It attempts to improve the player. It also creates evidence capable of changing the model. The frozen periodization sequence is $$ \text{BASE} \rightarrow \text{BUILD} \rightarrow \text{SHARPEN} \rightarrow \text{TAPER} \rightarrow \text{COMPETITION} \rightarrow \text{RECOVERY}. \tag{25} $$ Taper begins fourteen days before the campaign target and Competition begins two days before it in the frozen release. These are intervention-policy choices, not established universal laws of chess preparation. --- # 5. THE RESEARCH PLANE The Research Plane begins with the same evidence history but not the same estimators. A binary correctness result is not pooled with an ordinal process score simply because both can be represented numerically. A bounded operational encoding is not assumed to become interval measurement by normalization. Operational context weights and pseudo-priors do not enter a primary research contrast merely because the production scheduler already uses them. ## 5.1 The legacy operational Transfer Gap For dimension \(d\), the frozen implementation divides evidence into a supported pool containing `guided` and `curriculum` events and a transfer pool containing `mixed`, `otb`, `sparring`, and `serious-game` events. It computes $$ G_d^{\mathrm{op}} = \mu_d^{G} - \mu_d^{T}. \tag{26} $$ The shipped policy marks fewer than three raw transfer events as insufficient; gaps greater than \(0.18\) as material; gaps greater than \(0.08\) as watch conditions; and smaller gaps as the operational no-warning state labeled `stable` in the implementation (Office of Method, 2026d). The operational label **stable** does not establish scientific transfer stability. The quantity does not estimate transfer loss. ## 5.2 Proposition 1 — unequal-evidence shrinkage From Equation (23), $$ \mu_a(W;p) = p+ \frac{a(1-2p)}{2a+W}. \tag{27} $$ For two evidence pools \(G\) and \(T\), $$ G^{\mathrm{op}} = (p_G-p_T) + \left[ \frac{a(1-2p_G)}{2a+W_G} - \frac{a(1-2p_T)}{2a+W_T} \right]. \tag{28} $$ Equation (28) decomposes the operational difference into two components: $$ \text{observed-score difference} + \text{differential shrinkage}. $$ If both contexts have the same weighted observed score \(p\), $$ p_G=p_T=p, $$ then $$ \mu_a(W_G;p) - \mu_a(W_T;p) = \frac{ a(W_G-W_T)(2p-1) }{ (W_G+2a)(W_T+2a) }. \tag{29} $$ Therefore, for \(a>0\), \(p\neq0.5\), and \(W_G\neq W_T\), the operational context difference is generally non-zero even though observed performance is identical. This yields the central formal result. > **Proposition 1 — Unequal-evidence shrinkage.** Under the symmetric pseudo-prior accumulator used by the frozen operational model, unequal evidence mass can produce a non-zero supported-versus-transfer difference even when the weighted observed bounded performance score is identical in both contexts. Differentiating Equation (23) gives the equivalent monotonicity result $$ \frac{\partial\mu_a}{\partial W} = \frac{a(2p-1)} {(2a+W)^2}. \tag{30} $$ If \(p>0.5\), increasing evidence mass raises the shrunken mean toward \(p\); if \(p<0.5\), increasing evidence mass lowers it toward \(p\); if \(p=0.5\), evidence mass has no effect on the mean. For the frozen prior \((2,2)\), twenty perfect curriculum observations weighted at \(0.62\) produce $$ W_G=12.4 $$ and $$ \mu^G = \frac{14.4}{16.4} \approx0.878. $$ Three perfect serious-game observations produce $$ W_T=3 $$ and $$ \mu^T = \frac{5}{7} \approx0.714. $$ Thus $$ G^{\mathrm{op}} \approx0.164 $$ despite identical perfect observed performance. The defect is structural for the interpretation, not an implementation bug. The accumulator is performing the shrinkage it was designed to perform. What fails is the stronger inference that the difference isolates transfer. ## 5.3 Research-plane context discrepancy When evidence is on a common declared dimension-specific scoring family, a descriptive context mean can be written $$ \bar{S}_{d,c} = \frac{1}{n_{d,c}} \sum_{i\in\mathcal{I}_{d,c}} S_i. \tag{31} $$ A shrinkage-free descriptive supported-versus-transfer discrepancy is then $$ \Delta_d^{R} = \bar{S}_{d,G} - \bar{S}_{d,T}. \tag{32} $$ Equation (32) does not identify a causal transfer effect. Task difficulty, exposure, fatigue, source composition, occasion, modality, and skill mixture can still differ across contexts. Its narrower advantage is that the point estimate is not mechanically altered by pseudo-prior shrinkage or operational context coefficients. The Research Plane's primary default is therefore unweighted analysis within deliberately controlled or stratified comparisons. Design weights are admissible where they follow from known sampling probabilities. Operational importance weights are not imported automatically. ## 5.4 Weighted evidence mass is not independent sample size The operational quantity $$ W_d = \sum_{i\in\mathcal{I}_d} w_i \tag{33} $$ is called **weighted evidence mass**. It is not effective sample size. Ten related positions from one source do not become ten independent observations because their scalar weights sum to ten. Research reports should preserve raw count, evidence mass, source diversity, repeated-exposure status, occasion, and game clustering separately. ## 5.5 Item difficulty and identifiability The frozen Player Model does not contain a validated item-difficulty scale. Adding one item parameter to a future statistical model would not solve the problem automatically. In an N-of-1 longitudinal system, the player changes with time while exposure changes the item. If one evolving player sees each item once, player state and item difficulty are generally not separately identified without additional structure. Repeated presentation of the identical item is not a clean anchor for first-exposure ability because the repetition itself changes the learner. More defensible future approaches include parallel-form families, externally calibrated item pools, multiple participants, blinded expert difficulty ratings, deliberately crossed item/occasion designs, or appropriately constrained hierarchical models. Engine features such as branching, evaluation margin, tactical volatility, and WDL sensitivity can describe computational properties of a position. They are useful covariates. They are not automatically measures of human difficulty. ## 5.6 Item-to-skill mapping The Q-matrix formalism establishes a traceable record for dimension attribution. It does not imply that a binary Q-matrix is the final cognitive ontology. Each research-use item should identify: * primary modeled dimension; * secondary substantive dimensions; * dimensions that are merely incidental; * mapping-codebook version; * mapping source; * mapping confidence; * mapper identity. Q-matrix specification is itself a validity problem in cognitive-diagnostic research (Wang et al., 2026). CAPTUREDMIRAGE should treat its mapping procedure accordingly. Some chess tasks may be approximately conjunctive: catastrophic board-representation failure can make later calculation uninterpretable. Other tasks permit compensation. The present paper therefore does not introduce an unidentified multidimensional item-response model. Future validation will require deliberately crossed evidence capable of separating item, dimension, occasion, support, modality, and game effects. ## 5.7 Two distinct prospective questions The known defect in \(G_d^{\mathrm{op}}\) does not make the shipped warning scientifically uninteresting. It creates two different questions. The first is the **prospective field criterion profile of the operational warning**: > For each PRIMARY field-assessable dimension, what warning state was frozen before competition, and what serious-game SUCCESS, FAILURE, and INDETERMINATE profile was subsequently observed under the dimension-specific frozen criterion? This evaluates the warning as a practical pre-event signal while leaving its algebraic defect intact. The dimensions do not share a common outcome scale or known common base rate, so Field Event A defines no cross-dimensional concordance coefficient, discrimination statistic, or success threshold. Differences among dimension-specific failure proportions are descriptive. The second is **Research-Plane context discrepancy**: > After operational weights and pseudo-priors are removed, and only sufficiently comparable evidence is retained, does a meaningful supported-versus-transfer difference remain? Question 1 is the primary prospective field question. Question 2 is exploratory at Field Event A because the available pre-event evidence may not support a sufficiently controlled set of matched context comparisons. The two questions are not merged. --- # 6. PROCESS EVIDENCE, CHESS REFERENCE, AND SEMANTIC TRUTH ## 6.1 Calculation Forge as process localization Calculation Forge 5 organizes human search through $$ \text{ORIENT} \rightarrow \text{GENERATE} \rightarrow \text{RESIST} \rightarrow \text{CONCLUDE} \rightarrow \text{COMPARE} \rightarrow \text{CHECK} \rightarrow \text{COMMIT}. \tag{34} $$ The runtime can record candidate discovery order, analysis order, opponent responses, branch endpoints, branch conclusions, endpoint comparison, move-order checking, latency, final commitment, and confidence. The frozen calculation corpus contains 240 positions and a larger 24-class operational failure taxonomy. The purpose of process evidence is not to replace chess outcome. Tournament chess records the move that was played. Process evidence asks a different question: what divergence, if any, is supported by the available process record? Suppose available evidence supports two different reconstructions. One player selects the best move after considering only a cooperative opponent reply. Another identifies the relevant candidates and principal resistance accurately, reaches a difficult endpoint, mis-evaluates that endpoint, and selects the second-best move. The final-move metric favors the first player. A development system has a legitimate reason to treat the records differently. The qualification **if the evidence supports these reconstructions** is essential. Process classification is inference from imperfect records. It is not direct access to cognition. ## 6.2 Primary research process codebook The prospective field protocol uses a deliberately narrower research codebook than the operational 24-class taxonomy. The primary research classes are defined normatively in Appendix F and include: * representation divergence; * candidate omission; * opponent-resistance omission; * branch-control/termination failure; * endpoint-evaluation failure; * comparison/selection divergence; * move-order failure; * final recheck/verification failure; * multiple supported divergences with ordering indeterminate; * process indeterminate. The codebook does not force every bad move into a causal category. Indeterminate is a valid result. ## 6.3 The absence-of-mention rule Retrospective silence is not evidence that a thought never occurred. If a post-game record does not mention candidate \(m\), two explanations remain possible: the candidate was not considered, or the candidate was considered but was not remembered or reported later. Candidate-omission and opponent-resistance-omission codes therefore require a process record whose completeness is sufficient to support exclusion. Every primary retrospective process record receives a completeness classification: $$ C_i^{\mathrm{recall}} \in \{ \text{complete-enough}, \text{partial}, \text{fragmentary}, \text{unknown} \}. \tag{35} $$ Only `complete-enough` records can support a negative inference of candidate or resource omission. Silence in a partial, fragmentary, or unknown-completeness account is indeterminate. ## 6.4 Candidate equivalence A human candidate is not considered omitted merely because it differs from engine top-1. The primary chess reference distinguishes tablebase-eligible positions, forced-mate positions, and ordinary engine-evaluated positions. For ordinary positions, the acceptable candidate set is defined by the frozen engine protocol in Section 9.10. The important methodological requirement is **candidate completeness relative to the human record**. Initial MultiPV output does not define the full universe of acceptable human candidates. Every explicitly remembered human candidate absent from the initial MultiPV set receives a separate fixed-budget forced-root evaluation before candidate omission can be adjudicated. The purpose is not to declare Stockfish's preferred move the only humanly legitimate decision. It is to prevent the reference standard from excluding a human candidate merely because the engine was instructed to print too few root lines. ## 6.5 Process-report provenance Protocol-analysis research distinguishes concurrent reporting from retrospective reconstruction and emphasizes both reactivity and incompleteness (Ericsson & Simon, 1993). CAPTUREDMIRAGE therefore records how process evidence was obtained. Relevant provenance classes include: * concurrent interaction record; * timestamp-derived interaction record; * immediate pre-analysis retrospective reconstruction; * delayed retrospective reconstruction; * later analytical inference. These are not treated as epistemically interchangeable. Tournament process evidence is necessarily retrospective under the field protocol because research notation is not created during play. The primary field process record is therefore described as a **pre-analysis retrospective reconstruction**, not as direct observation of cognition. Contamination is broader than engine exposure. Discussion with the opponent, another player, a coach, a database, a book, an organizer analysis session, or any explanatory source can alter the reconstruction before an engine is ever opened. Primary field-process evidence therefore requires that the first-pass reconstruction occur before any chess analysis or analytical discussion. Where that condition cannot be met, the game remains available for objective move analysis, but no primary process inference is claimed from the contaminated reconstruction. --- # 7. REPAIR, CONTAMINATION, RETRIEVAL, AND PHYSICAL TRANSFER Retrieval practice can improve later retention (Roediger & Karpicke, 2006; Karpicke & Roediger, 2007). That benefit creates a measurement complication for a system that also wants independent evidence: the act of testing changes the learner and changes the evidentiary status of the item. CAPTUREDMIRAGE therefore treats exposure as provenance. Instructional success cannot simultaneously serve as independent proof that the instruction generalized. ## 7.1 The repair loop The closed loop is $$ \text{PLAY} \rightarrow \text{RECONSTRUCT} \rightarrow \text{DIAGNOSE} \rightarrow \text{REPAIR} \rightarrow \text{RETEST} \rightarrow \text{TRANSFER} \rightarrow \text{UPDATE} \rightarrow \text{PLAY}. \tag{36} $$ A serious-game failure may generate a repair case. Instruction and near-transfer exercises can establish that the correction has been understood. Repair closure requires later evidence under reduced support and an independently reserved form. A failed sealed retest consumes the form but does not certify the repair. ## 7.2 Independence profile A different FEN is not automatically an independent test. Two positions can share a source game, move neighborhood, solution mechanism, thematic cue, or prior exposure. The research doctrine therefore records five forms of independence: **source independence** — no immediate common source record; **move-neighborhood independence** — not a nearby state in the same game sequence; **solution independence** — not a trivial transposition or memorized continuation of the repaired solution; **thematic-cue independence** — the relevant concept is not announced by category or presentation; **exposure independence** — the item or solution has not previously been revealed. For a sealed repair test to count as primary independent evidence, exposure, move-neighborhood, solution, and thematic-cue independence are mandatory. Source independence is also required unless a domain-specific exception was frozen before the item was encountered. Exceptions are recorded. They are not created after performance. ## 7.3 Consumption and benchmark epochs Exposure changes assessment status: $$ \text{sealed} \longrightarrow \text{consumed}. \tag{37} $$ Consumption does not erase the item. It changes what the item can measure. A consumed position can remain useful for retrieval, reconstruction, or instruction. It cannot return to the first-exposure pool. Benchmark replenishment therefore creates a longitudinal measurement problem. New epochs should preserve form composition, item provenance, exposure history, approximate complexity descriptors, linking methodology where available, and the reason for epoch change. Scores from different benchmark epochs are not assumed equivalent merely because both are called Benchmark Core. ## 7.4 Physical-board transfer The expertise literature establishes the importance of structured domain-specific representation, but it does not establish that this particular participant will exhibit a meaningful physical-versus-digital difference. CAPTUREDMIRAGE therefore treats modality as an evidence facet and the modality effect as a hypothesis. OTB Companion excludes board setup from decision time, makes the physical board authoritative after sealing, and records a digital reveal as an evidence-context downgrade. In physical Tournament Simulation, synchronization of an opponent move onto the physical board does not consume the player's simulated decision clock. These are measurement controls. Experimental overhead should not be mistaken for chess decision time. ## 7.5 Modality crossover A future within-player modality study should use parallel rather than identical forms. Items should be randomized to physical or digital presentation, matched-pair identities should be concealed from the participant where practicable, and the second condition should use different but comparable material after adequate delay. Potential outcomes include reconstruction error, candidate latency, move quality, decision time, confidence, and process-localization category. Matching variables can include material, structural family, tacticality, game phase, and computational complexity descriptors without assuming that those descriptors constitute validated human difficulty. --- # 8. ADAPTIVE TRAINING AND RELATED WORK The Player Model is primarily formative. Its purpose is not to replace chess rating with a more elaborate global score. It asks where the next marginal unit of training or measurement should go. Knowledge-tracing systems have modeled changing learner state for decades (Corbett & Anderson, 1995), while cognitive-diagnostic models explicitly represent multidimensional attribute profiles rather than one scalar ability (Wang et al., 2026). The architecture examined here is therefore not novel because it decomposes performance into dimensions. Its focus is the integration of evidence provenance, support withdrawal, contamination state, process localization, physical modality, sealed retesting, and serious-game linkage inside one longitudinal control system. The deliberate-practice literature provides another constraint. Ericsson et al. (1993) developed the influential deliberate-practice framework; Charness et al. (2005) found serious chess study strongly associated with chess expertise in their samples; Macnamara et al. (2014) showed that deliberate practice does not explain all performance variance. Southwick et al. (2026), using objective longitudinal Chess.com data from 44,213 players, found substantial differences in improvement per unit time among forms of chess activity. That work concerns online behavior and online rating change. It does not establish over-the-board transfer, validate CAPTUREDMIRAGE's scheduler, or imply that the most efficient online practice category is necessarily the most efficient route to physical serious-game generalization. CAPTUREDMIRAGE therefore asks a narrower unresolved question: > **Can evidence-conditioned allocation improve the efficiency or transfer value of structured practice relative to simpler allocation policies?** ## 8.1 Single-case experiments A single participant can support rigorous experiments on narrower components, but longitudinal observation alone is not an experimental design. Single-case experimental methodology requires repeated measurement and deliberate manipulation of conditions (Krasny-Pacini & Evans, 2018). Chess learning also creates irreversibility. Once a representation or technique has been learned, withdrawal cannot restore the original state cleanly. Multiple-baseline designs may be possible across sufficiently distinct skill families. Alternating-treatment designs are appropriate only where carryover is limited enough that one condition does not permanently alter the next. Randomized parallel-form modality studies and staggered introduction of repair procedures are therefore more defensible for some questions than simple A/B alternation. ## 8.2 Stage progression The six operational mastery stages can eventually be evaluated prospectively by testing whether passage under one evidence condition predicts performance under the next relevant condition while controlling task composition and exposure as far as feasible. Weak prospective discrimination would require recalibration or later modification of the stage architecture. ## 8.3 Testing whether provenance adds information One of the architecture's central practical premises is itself testable by ablation. From the same frozen longitudinal ledger, construct: **Full representation:** item, native result, modeled dimension, support, cue state, exposure, process provenance, modality, source, occasion, and serious-game context. **Outcome-only representation:** item identity, time, and raw performance. The empirical question is whether the full representation improves later prediction, calibration, diagnosis, or training allocation relative to the outcome-only representation. If it does not, an important practical justification for the added provenance architecture is weakened. This paper makes no priority claim for the integrated system. --- # 9. PROSPECTIVE FIELD EVALUATION ## 9.1 Public and embargoed event identity The prospective study concerns **Field Event A**, a six-round, in-person classical Swiss tournament under a classical time control with increment. The author is the participant. The public manuscript deliberately withholds the event name, venue, exact dates, exact schedule, exact time control, exact section, federation identifiers, rating-administrative details, and direct registration record. Those particulars are unnecessary to evaluate the scientific protocol before the event and would create an avoidable correlation bridge between the public pen name and external chess records. The exact event identity, registration evidence, schedule, time control, section, applicable rating/rules status, and participant administrative identifiers are retained in the embargoed Pre-Event Lock. This is not participant concealment. The study remains explicitly disclosed as an author-participant self-study. The distinction is between **scientifically necessary self-study disclosure** and **unnecessary publication of a cross-identity lookup key**. ## 9.2 Research designation Field Event A is > **a prospective single-participant ecological field evaluation and protocol stress test.** It is not an efficacy trial. There is one participant, six games, variable opponents, variable openings, non-random tournament pairings, no concurrent control condition, and only one event. Decisions are nested within games, and games are nested within the same tournament environment. Rating change cannot establish treatment efficacy. The event has three narrower purposes: 1. expose the frozen evidence architecture to serious-game observations; 2. construct a prospective warning–criterion profile across the frozen PRIMARY field-assessable dimensions; 3. stress-test the sampling, mapping, coding, and reference procedures before a larger longitudinal tournament series. ## 9.3 Author-participant role The author designed CAPTUREDMIRAGE, operates the system, and is the participant in the prospective evaluation. This creates intellectual-allegiance, expectancy, intervention-awareness, self-report, recognition, and coding risks. The design responds by separating operational warnings from research estimands, freezing selection and mapping procedures before outcome analysis, using sequential information masking where possible, preserving indeterminate cases, exposing the no-alert/no-opportunity states in advance, and externally timestamping the Pre-Event Lock. No participant-level or rater-level blinding is claimed when the author performs a research step. Masking designated data fields cannot make the author unaware of a game he played, and he may recognize positions even when the played move is withheld. The operational system may already have exposed the participant to some of its weakness estimates during ordinary training. The prospective test therefore concerns the future field criterion profile of a shipped operational warning under an alert-aware training regime. It is not a first-ever transfer from laboratory training into serious play. ## 9.4 Tournament legality No research notation is created during play. The participant records only information permitted by the governing tournament rules. No candidate codes, confidence marks, uncertainty flags, research annotations, devices, databases, engines, or external analysis are used during the game. The final governing rule documents and event-specific organizer notes are included in the embargoed lock. Where national and international rules overlap, the research protocol adopts the stricter applicable prohibition on analytical notes and outside information. ## 9.5 Two post-game corpora The field study separates diagnostic mining from prospective evaluation. ### Corpus A — Diagnostic Critical Positions Corpus A is intentionally outcome-enriched. After the game and after the primary pre-analysis record has been frozen, it may include positions selected because of: * large engine-evaluation loss; * material WDL deterioration; * tactical turning point; * severe time expenditure where reliable timing exists; * strategically consequential decision; * recurrence of a known repair class. Corpus A exists for diagnosis and later repair. It is not used to estimate process-failure prevalence. ### Corpus B — Outcome-Independent Evaluation Sample Corpus B is selected without using the quality of the participant's eventual move. Eligible decisions satisfy all of the following: 1. the position belongs to a played tournament game rather than a bye, forfeit, or administrative non-game; 2. the participant is to move; 3. the fullmove number is at least 8; 4. at least two legal moves exist; 5. the position can be reconstructed legally from the game score. The fullmove-8 cutoff is a deliberate design choice. It reduces domination of the generic evaluation corpus by early memorized opening retrieval. Opening-specific dimensions are not thereby treated as absent from the event; they require separate criteria under the Field Criterion Map rather than being forced into Corpus B. Eligible decisions are partitioned into pragmatic **temporal strata**: * Stratum A: fullmoves 8–18; * Stratum B: fullmoves 19–35; * Stratum C: fullmoves 36 and later. These are not labeled opening, middlegame, and endgame. Up to two positions are selected from each stratum in each game. If a stratum contains fewer than two eligible positions, all eligible positions in that stratum are retained and the deficit is not backfilled. Selection is deterministic. For each eligible position construct `OM04-WP001-FIELD-A-B2|round=|ply=

` as an exact UTF-8 byte string with no trailing newline, where `` is the round number and `

` is the ply index immediately before the participant's move. Compute SHA-256. Within each game/stratum, select the two lexicographically smallest hexadecimal digests. FEN is deliberately excluded from the ranking payload so that different valid FEN serialization conventions cannot alter sample membership. The exact FEN remains stored in the resulting sample manifest. This is an **outcome-independent deterministic sample**. It is not described as a probability sample, an unbiased sample, or a representative sample of all tournament decisions. ## 9.6 Pre-analysis reconstruction and measurement burden The strongest process evidence available in ordinary tournament conditions is still retrospective. After each game, before any chess analysis or analytical discussion, the participant creates a rapid chronological first-pass reconstruction. The legal game score may be used to recover the move sequence, but no engine, database, coach, opponent post-mortem, organizer analysis session, book, or explanatory source is consulted. The record captures remembered decision moments and permits explicit uncertainty. It does not require the participant to invent a thought account for every move. For every described decision, the participant records recall completeness as: * complete-enough for omission inference; * partial; * fragmentary; * unknown. For an endpoint judgment to enter the PRIMARY endpoint-evaluation criterion, the participant must also have recorded one of three explicit root-side labels before analysis: * **favorable**; * **approximately balanced**; * **unfavorable**. The process-measurement procedure is intentionally bounded so that the measurement system does not consume a disproportionate share of tournament recovery. For a game followed by at least 60 minutes before the next scheduled round, the first-pass reconstruction is capped at **15 minutes**. If fewer than 60 but at least 30 minutes remain before the next scheduled round, it is capped at **8 minutes**. If fewer than 30 minutes remain, no primary process reconstruction is required. The following times are recorded: * game end; * reconstruction start; * reconstruction end; * reconstruction duration; * scheduled start of the next round. The game remains in the objective move corpus even when no primary process record can be obtained. Primary process evidence is **not assumed missing at random**. Long, difficult, emotionally consequential, or late-ending games may be more likely to lack complete reconstruction. Process-complete games are therefore not treated as an unbiased subset of all tournament games. Only after the first-pass reconstruction is frozen is Corpus B generated. A later position-cued reconstruction can be collected for diagnosis, but it is classified as secondary prompted retrospective evidence and excluded from primary process inference. ## 9.7 Event-bounded target domain and task-demand mapping For dimension \(d\), the event-bounded field domain is $$ \mathcal{G}^{A}_d = \{ \text{Corpus B decisions classified as field-assessable opportunities for }d \}. \tag{38} $$ The actual tournament positions do not exist before the event, so their Q-matrix rows cannot be preregistered. What is frozen before the event is: * the skill taxonomy; * the task-demand mapping codebook; * primary/secondary dimension definitions; * mapping-confidence rules; * field-assessability rules; * mapper procedure. After Corpus B has been selected, an **outcome-field-masked mapping packet** is generated containing only: * sampled position identity; * FEN; * side to move; * complete legal-move list. The packet does not contain: * the participant's move; * pre-event warning status; * process reconstruction; * engine evaluation of the participant's move; * eventual process code. For the four PRIMARY dimensions defined below, opportunity status must be assignable from these fields together with the frozen dimension definitions. If a later criterion requires additional reference information merely to determine opportunity status, that dimension is not promoted into the PRIMARY analysis for Field Event A. The mapper assigns dimension relevance before the withheld outcome fields are revealed. If the author performs the mapping, the procedure provides **sequential information masking**, not blinding: the mapping artifact is completed and hashed before the human move, warning state, process record, and outcome adjudication are opened for failure coding. No claim of participant-level or rater-level blindness is made. If an independent mapper is available, that fact and the fields available to that mapper are recorded separately. The resulting event Q rows are generated post-event under the frozen procedure and then versioned and hashed. ## 9.8 Field Criterion Map and primary adjudication frame Not every one of the 43 Player Model dimensions can be assessed defensibly from six tournament games. The complete field-assessability classification appears in Appendix H. The PRIMARY prospective profile for Field Event A is deliberately restricted to four calculation dimensions for which an executable SUCCESS/FAILURE/INDETERMINATE rule can be stated without inventing a post-event strategic judgment threshold: $$ \mathcal{F} = \{ \text{candidate-generation}, \text{opponent-resistance}, \text{endpoint-evaluation}, \text{comparison} \}. \tag{39} $$ `branch-control` and `move-order` remain operational dimensions and can produce secondary descriptive evidence, but they are classified as CONDITIONAL for Field Event A because a sufficiently general, non-arbitrary primary quantitative criterion is not frozen here. The PRIMARY adjudication rules are as follows. ### Table 5. Primary Field Adjudication Table | Dimension | Opportunity condition | Minimum evidence for scoring | SUCCESS | FAILURE | INDETERMINATE | Objective/reference rule | |---|---|---|---|---|---|---| | candidate-generation | Corpus B position mapped candidate-generation relevant | Recall record `complete-enough`; every explicitly recalled candidate legally identified; reference stable | At least one explicitly recalled candidate belongs to the frozen acceptable-candidate set | No explicitly recalled candidate belongs to the frozen acceptable-candidate set | Recall not complete-enough; candidate identity unresolved; or reference indeterminate | Candidate acceptability under Sections 9.10–9.13 | | opponent-resistance | A recalled candidate branch contains at least two legal opponent replies and is mapped opponent-resistance relevant | Branch record `complete-enough`; recalled opponent replies legally identified; reply reference stable | At least one explicitly recalled opponent reply belongs to the principal-resistance set | No explicitly recalled opponent reply belongs to the principal-resistance set | Branch recall not complete-enough; reply identity unresolved; or reference indeterminate | At the opponent node, principal-resistance set = replies within both 50 cp and 0.05 root-side expected-score loss of the best defensive reply under symmetric adjudication; tablebase/mate rules applied analogously | | endpoint-evaluation | A recalled branch reaches a legally reconstructable endpoint and contains an explicit pre-analysis favorable/balanced/unfavorable label | Endpoint FEN reconstructable; label recorded before analysis; reference stable | Human label equals objective three-way band | Human label differs from objective band by two bands: favorable vs unfavorable or unfavorable vs favorable | Adjacent-band disagreement; missing label; endpoint not reconstructable; or reference indeterminate | Objective band: favorable if root-side \(E_{\mathrm{WDL}}\ge0.60\); balanced if \(0.400\), $$ R_d^B = \frac{F_d^B}{C_d^B}. \tag{44} $$ Every report shows $$ N_d^B,\quad I_d^B,\quad C_d^B,\quad F_d^B,\quad R_d^B. $$ Indeterminate cases therefore cannot enter the denominator as hidden successes. One sampled position can be relevant to several dimensions. Skill-level denominators overlap and are correlated. They are not summed as though the categories were mutually exclusive. The proportions describe the sampled Field Event A evaluation corpus. They are not interpreted as unbiased estimates of failure prevalence across every decision in the tournament, much less across serious chess generally. ## 9.10 Candidate-complete and budget-symmetric engine-analysis protocol The primary objective reference uses Stockfish 18 under an exact Pre-Event Lock configuration. Stockfish's UCI interface distinguishes MultiPV output from `searchmoves` root restriction; the protocol therefore uses MultiPV for discovery and single-root searches for adjudication (Stockfish Developers, 2026). The lock records: * exact Stockfish executable and SHA-256; * executable architecture/build flavor; * operating-system name and version; * CPU model; * complete UCI option dump; * startup identification output; * filename and SHA-256 of every neural-network file reported as loaded by the exact executable, including primary and small networks where exposed; * Threads = 1; * Hash = 512 MB; * discovery MultiPV = 5; * `UCI_ShowWDL = true`; * `UCI_Chess960 = false`; * no strength limiting; * primary budget = 10,000,000 nodes per adjudicated root move; * output-parser identity; * exact tablebase configuration; * exact command sequence. Before each independent search, the engine is reset using the frozen sequence: 1. set the frozen options; 2. `ucinewgame`; 3. `isready`; 4. wait for `readyok`; 5. `position fen `; 6. issue the applicable `go` command. ### Stage A — discovery Run one 10,000,000-node MultiPV-5 search. This search identifies a discovery set only. Its evaluations are not used directly in the final candidate-difference calculation. ### Stage B — symmetric root adjudication Create the adjudication set from: 1. all five Stage-A root moves returned when five exist; 2. every explicitly recalled human candidate not already in that set; 3. any additional root move required by a frozen primary process criterion. Each adjudication-set move receives its own reset, single-root `searchmoves ` search at exactly 10,000,000 nodes. The **reference best move** for the primary classification is the move with the best root-side result among this symmetrically adjudicated set. The paper does not claim that this finite set exhaustively proves the globally best legal move; it defines a reproducible candidate reference sufficient for the frozen field criteria. The same budget symmetry applies at the 50,000,000-node stability escalation. ### Opponent-resistance adjudication At a recalled candidate branch's opponent node, use the same two-stage procedure: MultiPV-5 discovery followed by separate equal-budget root searches for each discovery reply and each explicitly recalled opponent reply. The best defensive reply and principal-resistance set are determined from these symmetric searches. ## 9.11 WDL expected score and ordinary candidate equivalence When Stockfish returns WDL values in permille, define root-side expected score as $$ E_{\mathrm{WDL}} = \frac{W+\tfrac12D}{1000}. \tag{45} $$ For candidate \(m\), $$ \Delta E(m) = E_{\mathrm{best}}-E_m, \tag{46} $$ with all values normalized to the root side to move. For ordinary non-mate, non-tablebase positions, a candidate belongs to the primary acceptable set only when both $$ \Delta \mathrm{CP}(m) \le 50 \tag{47} $$ and $$ \Delta E(m) \le0.05. \tag{48} $$ These are protocol thresholds. They are not asserted as universal definitions of a good human move. ## 9.12 Tablebase and forced-mate candidate rules ### Tablebase-eligible positions Where the locked Syzygy installation establishes an exact WDL result under the frozen 50-move-rule convention, primary candidate acceptability is based on **preservation of the root-side WDL class**. A move is acceptable if it preserves the exact root-side WDL class of the reference-best tablebase move. A move that changes win to draw/loss, draw to loss, or otherwise worsens the exact root-side WDL class is unacceptable. DTZ is recorded descriptively but does not alter the PRIMARY candidate-generation classification in Version 1.0. Cursed wins and blessed losses are interpreted under the locked 50-move-rule setting and the tablebase WDL semantics recorded in the Pre-Event Lock. ### Forced-mate positions Where the symmetric locked engine reference establishes a forced mate for the root side, a candidate is primary-acceptable if its symmetric adjudication also preserves a forced mate for the root side. Mate distance is recorded descriptively but does not determine the PRIMARY candidate-generation category. If forced-mate status changes between the 10M and required 50M stability analyses, the relevant reference classification is **ENGINE-REFERENCE INDETERMINATE**. ## 9.13 Reference-stability gate A fixed search budget can be reproducible without being substantively stable. The protocol therefore allows the reference instrument to return **ENGINE-REFERENCE INDETERMINATE**. A mandatory 50,000,000-node symmetric escalation is performed when any primary candidate classification lies in the boundary band $$ 35 \le \Delta \mathrm{CP} \le 65 \tag{49} $$ or $$ 0.035 \le \Delta E \le 0.065. \tag{50} $$ The same escalation occurs if the parser identifies mate-score instability or a discrepancy between exact tablebase evidence and an ordinary engine classification. Every adjudication-set move material to the boundary decision is escalated under the same single-root budget. If the 10M and 50M analyses fall on different sides of the primary acceptable/unacceptable boundary, the engine reference is coded indeterminate for that classification rather than selecting whichever depth produces a preferred research result. ## 9.14 Syzygy boundary Where tablebases are available and the position is eligible, the lock records: * exact Syzygy file-set manifest and hashes; * maximum supported piece count; * `Syzygy50MoveRule`; * `SyzygyProbeLimit`; * `SyzygyProbeDepth`; * treatment of cursed wins and blessed losses; * parser behavior for WDL and DTZ. The intended primary configuration uses the ordinary 50-move-rule convention. No position is described as tablebase-exact unless the required files are present and verified under the locked configuration. ## 9.15 Primary prospective question: warning–criterion profile The shipped operational warning is evaluated as a warning. It is not relabeled as a transfer estimator. The final research snapshot is taken only **after the participant's final substantive pre-event CAPTUREDMIRAGE training session has ended** and before the first round. The exact ISO 8601 timestamp is recorded in the embargoed lock. No scored CAPTUREDMIRAGE training, targeted repair, or alert-specific remediation occurs between the snapshot and the first round. If such an intervention does occur, the affected dimension is marked $$ \text{INTERVENED-BEFORE-CRITERION} $$ and is not treated as a clean pre-event warning observation. For every dimension \(d\), define the locked operational alert indicator $$ A_d = \begin{cases} 1,&\text{if the shipped transfer-gap rule reports a material warning},\\ 0,&\text{otherwise}. \end{cases} \tag{51} $$ The shipped rule requires at least three raw transfer-context events and $$ G_d^{\mathrm{op}}>0.18. $$ For every alert, the report also publishes: * raw supported-context count; * supported weighted evidence mass; * raw transfer-context count; * transfer weighted evidence mass; * supported operational mean; * transfer operational mean; * operational gap. The primary field report shows \(A_d\) and the full opportunity/outcome record for **every** dimension in the PRIMARY frame \(\mathcal{F}\), not only dimensions that happened to generate warnings. The reported object is $$ (A_d,N_d^B,I_d^B,C_d^B,F_d^B,R_d^B) $$ for all \(d\in\mathcal{F}\), together with SUCCESS counts derivable as \(C_d^B-F_d^B\). No single cross-dimensional concordance coefficient, discrimination statistic, accuracy statistic, or success threshold is defined for Field Event A. The four criteria do not have known common base rates or a common measurement scale, so cross-dimensional differences in \(R_d^B\) are descriptive. If no dimension in \(\mathcal{F}\) is operationally alerted at the locked snapshot, Question 1 is reported as **no analyzable alert set**. If all dimensions in \(\mathcal{F}\) share the same alert state, no alert-versus-non-alert comparison exists. No replacement hypothesis is chosen after the event. If an alerted dimension receives no classifiable field opportunity, that dimension contributes **no criterion result**, not a success. The field profile can supply preliminary dimension-specific criterion evidence concerning the warning's practical signal. It does not validate transfer measurement. ## 9.16 Exploratory Research-Plane context discrepancy Question 2 remains exploratory at Field Event A. Where a dimension has a common scoring family across supported and transfer evidence and the pre-event evidence extract permits a defensible comparison, Equation (32) is reported together with the underlying task composition. No pseudo-prior or operational context coefficient enters the descriptive point estimate. Because task comparability cannot yet be reduced to a fully validated mechanical criterion across all dimensions, this analysis is not promoted to a confirmatory field endpoint. ## 9.17 Repair recurrence Repairs completed before the field event are reported separately. A repair can enter the recurrence analysis only when its frozen Field Criterion Map supplies an observable opportunity/failure rule. For each repair-linked dimension or failure family, the report shows: * relevant opportunities; * indeterminate opportunities; * classifiable opportunities; * recurrences. Zero failures with zero relevant opportunities is no test. Zero failures across several relevant classifiable opportunities is favorable descriptive evidence, not proof of permanent repair. ## 9.18 Between-round adaptation A perfectly frozen psychological state is impossible. Game 1 can alter Game 2 even without formal training. During Field Event A, all substantive chess activity between rounds is logged, including: * post-mortem analysis; * opening preparation; * organizer or coach analysis; * lectures with direct analytical relevance; * side-event play; * CAPTUREDMIRAGE use; * other substantive study. No new targeted CAPTUREDMIRAGE repair directed specifically at the locked PRIMARY alert set is intentionally introduced before completion of the event. If an intervention materially targets a PRIMARY dimension, subsequent observations for that dimension are marked post-intervention. The ordinary adaptive repair loop resumes after the event. ## 9.19 Inter-rater and intra-rater coding Independent second coding and delayed self-recoding answer different questions. If a qualified independent coder is secured before the Pre-Event Lock, **all primary classifiable cases submitted to the main process analysis are independently coded**. Any case unavailable to the second coder for a documented technical or procedural reason is identified individually and excluded from the complete-case inter-rater agreement calculation; it is not selectively omitted because of difficulty or disagreement. If no qualified independent coder is secured before the lock, the paper states that no independent inter-rater estimate is available. A delayed self-recode can still be performed after a minimum fourteen-day interval to examine intra-rater consistency, but that result is reported explicitly as intra-rater agreement and is not substituted for independent coding. Disagreements remain visible in the research record. ## 9.20 Six-game inferential boundary No tournament-level significance testing is applied. No random-effects game distribution is treated as well estimated from six clusters. No rating change is attributed causally to CAPTUREDMIRAGE. The field report includes: * game-by-game objective results; * outcome-independent Corpus B selections; * task-demand mappings; * process provenance; * completeness classes; * engine-reference status; * opportunity denominators; * failures; * successes; * indeterminate cases; * intervention boundaries; * protocol deviations; * leave-one-game-out descriptive sensitivity where meaningful. The event can supply prospective evidence about particular precommitted claims. It cannot establish population-level efficacy. ## 9.21 Pre-Event Lock The public protocol is complete before the event, but some event-dependent artifacts cannot exist until final preparation. The Pre-Event Lock instantiates those values without altering the governing rules. It contains: * exact event identity; * venue and dates; * exact schedule and time control; * exact section; * registration record; * applicable national/international rating and rule status; * participant administrative identifiers required for event verification; * final research-snapshot timestamp; * frozen Player Model export; * complete operational alert list; * full PRIMARY field-assessable frame; * Field Criterion Map version; * skill taxonomy; * item-to-skill mapping codebook; * process-codebook version; * Corpus B selector implementation; * mapping-packet generator; * field-analysis implementation; * exact Stockfish executable and SHA-256; * architecture/build flavor; * operating-system name/version and CPU model; * complete UCI option dump and startup identification output; * every neural-network file reported as loaded by the engine, with filename and SHA-256; * Syzygy manifest/configuration; * engine-output parser; * exact protocol/manuscript version; * deviation log initialized with any pre-event deviations. The lock does **not** contain nonexistent future Q rows for tournament positions. Those are generated after the event under the frozen mapping procedure. The lock uses `PRE_EVENT_LOCK_MANIFEST.json` as the artifact inventory. Each entry records: * normalized relative path; * byte size; * SHA-256; * semantic role. The manifest is serialized using the **JSON Canonicalization Scheme (JCS), RFC 8785**, and the exact canonical bytes are UTF-8 as required by that scheme (Rundgren, Jordan, & Erdtman, 2020). Define the human-readable lock identifier as $$ H_{\mathrm{lock}} = SHA256(\text{exact RFC-8785 canonical manifest bytes}). \tag{52} $$ OpenTimestamps is run **directly against the exact canonical `PRE_EVENT_LOCK_MANIFEST.json` file**, not against a separately created text representation of its hexadecimal digest (OpenTimestamps Project, 2026). The resulting `.ots` proof is retained with the lock. The lock records: * OpenTimestamps client/tool identity and version; * time of initial stamping operation; * initial proof status; * later proof-upgrade status where applicable; * later verification status where applicable. The separately reported \(H_{\mathrm{lock}}\) is the human-readable SHA-256 identifier of the same bytes whose file is timestamped. The timestamp proof and exact identifying contents can remain embargoed until their release no longer defeats the public author-identity compartment. Timestamping establishes pre-event existence of the committed bytes; it does not make a non-public artifact independently inspectable. Any post-lock deviation receives: * timestamp; * reason; * affected artifact or rule; * whether the deviation was known before outcome inspection; * effect on primary versus secondary status. The historical lock is not rewritten. --- # 10. CLAIMS, FAILURE MODES, LIMITATIONS, AND RESEARCH INTEGRITY ## 10.1 Four claim levels **Level I — software-system claims** concern implementation properties: build identity, evidence durability, sealed-pool behavior, state transitions, scheduler execution, and similar mechanisms. **Level II — measurement claims** concern whether modeled quantities behave as intended: process-code reproducibility, context discrepancy, stage progression, modality effects, or prospective warning–criterion profiles. **Level III — external-performance claims** concern serious-game behavior: recurrence, decision quality, tournament performance, rating development, or similar competitive outcomes. **Level IV — causal efficacy** concerns whether CAPTUREDMIRAGE itself produces superior outcomes to a relevant alternative. Evidence does not move automatically from one level to the next. ## 10.2 Failure taxonomy A **calibration failure** occurs when a construct remains useful but thresholds, priors, coefficients, or timing policies perform poorly. An **algorithmic or estimand failure** occurs when a formula cannot estimate the quantity assigned to it even under favorable data. The unequal-shrinkage problem of the operational Transfer Gap as a transfer estimator is an example. A **construct failure** occurs when a measure does not correspond adequately to the competency it is intended to represent. Poor process-code reproducibility or disappearance of a context discrepancy after task control would weaken the construct. An **intervention failure** occurs when a meaningful diagnosed weakness is not improved by the assigned repair procedure. An **external-performance failure** occurs when internal CAPTUREDMIRAGE metrics improve while serious-game behavior does not. A **causal-efficacy failure** concerns the strongest claim: use of CAPTUREDMIRAGE does not outperform a relevant alternative under a design capable of identifying that difference. These outcomes are not to be repaired rhetorically by changing the target after the result is known. ## 10.3 Present limitations The project has one principal participant. Item difficulty is not calibrated. The operational context field is compound. The item-to-skill mapping procedure is authored within the same project. The frozen Player Model uses heuristic weights and pseudo-priors. Process evidence is partly introspective, and tournament process evidence is retrospective. Retrospective silence is therefore not treated as evidence of omission without a completeness condition. Many positions implicate several dimensions. Benchmark replenishment can create longitudinal drift. The scheduler is heuristic. Operational prior and context coefficients are design parameters. Current software assurance is internal rather than independent third-party certification. The author knows the architecture and cannot be treated as naive to the intervention. Field-process data may be missing informatively because long or difficult games can make reconstruction less feasible. The reconstruction procedure itself imposes burden and may affect recovery; this is why it is capped explicitly. Actual Q rows for new tournament positions are necessarily generated after the event under a precommitted codebook rather than literally preregistered in advance. The field event contains only six games and cannot resolve most of these limitations. ## 10.4 Author role and competing interests The author designed and operates CAPTUREDMIRAGE and is the participant in the prospective evaluation. This creates intellectual-allegiance, expectancy, intervention-awareness, coding, and self-report risks. The research response is procedural rather than rhetorical: claims are restricted; negative and indeterminate results are preserved; task-demand mapping is separated from failure coding; pre-event artifacts are committed before outcome inspection; independent coding is distinguished from self-recoding; and no field result is promoted automatically to causal efficacy. The author has an intellectual interest in the success of CAPTUREDMIRAGE as an Office of Method research system. No external commercial sponsorship is represented in this paper. Any later material commercial interest should be disclosed in the version in which it exists. ## 10.5 Participant and ethical scope The present field work is an author-participant self-study. The paper analyzes the author's own training evidence, decision records, and retrospective process reports. Tournament games necessarily involve opponents, but opponent internal states are not treated as participant data. Later publication of game records should minimize unnecessary opponent-identifying information where doing so does not compromise the chess record required for analysis. No claim is made that an independent institutional review board approved this self-study. Expansion to recruited participants would require a separate protocol addressing informed consent, privacy, data handling, withdrawal, and applicable institutional ethical requirements. ## 10.6 Novelty and scope This paper does not claim invention of transfer, retrieval practice, knowledge tracing, cognitive diagnosis, deliberate practice, protocol analysis, chess expertise, physical representative practice, or adaptive scheduling. It does not make a priority claim for the integrated CAPTUREDMIRAGE architecture. Its present contribution is methodological and architectural: a longitudinal chess-development system designed so that instructional support, cue state, exposure, process provenance, semantic authority, modality, and serious-game evidence remain distinguishable before operational aggregation destroys those distinctions. Whether that added provenance improves prediction, diagnosis, allocation, or competitive development remains an empirical question. --- # CONCLUSION The central measurement problem in chess training is not that performance cannot be observed. It is that an observed performance rarely carries its own interpretation. A correct move in a named tactical exercise and a correct move found without thematic cues in a serious game are both real observations. They answer different questions. A repaired position solved immediately after instruction and the same mechanism recognized months later in an unrelated game are both useful evidence. They support claims of different breadth. A physical-board result, a sealed benchmark, and a guided strategic exercise differ along several evidentiary facets; no universal scalar hierarchy captures all of those differences. CAPTUREDMIRAGE is most defensible when understood as a provenance-preserving architecture before it is understood as a scoring model. Its operational system necessarily compresses evidence because it must decide what to train. That compression is useful, but it is not scientifically neutral. Its legacy context hierarchy mixes several dimensions. Its recency mechanism produces prior reversion rather than a validated forgetting law. Its beta-style accumulator uses heuristic pseudo-counts. Its operational Transfer Gap can be generated partly by unequal evidence mass even when weighted observed performance is identical. Those findings do not invalidate the operational system. They determine the limits of inference. The Research Plane therefore begins earlier. It preserves the native result, identifies the dimension actually being scored, retains process content separately from process provenance, requires task-demand mapping, refuses to infer omissions from retrospective silence, decomposes evidence contexts, distinguishes target distribution from sampling design, and excludes the operational pseudo-prior from transfer estimation. The prospective field study applies the same discipline to competition. Outcome-enriched diagnostic positions are separated from an outcome-independent evaluation sample. The public author remains explicitly the participant while unnecessary tournament identifiers are withheld in an embargoed pre-event commitment. The Player Model cannot define its own tournament failures after the tournament because the Field Criterion Map exists first. A candidate cannot be labeled omitted because the engine printed only five lines; every remembered candidate used in that judgment receives a reference search. An unavailable opportunity does not count as successful transfer. Missing process evidence remains indeterminate. A pre-event warning is not called predictive merely because one later failure can be narrated as relevant; the complete PRIMARY field-assessable profile is reported dimension by dimension. The system may still fail in several distinct ways. Operational warnings may show little useful correspondence with the dimension-specific serious-game criterion profiles. Process codes may prove unreliable. Sealed repairs may recur. Physical-board effects may be negligible. The scheduler may allocate training inefficiently. Added provenance may fail to improve prediction over an outcome-only representation. Internal metrics may improve while competitive behavior remains unchanged. Those possibilities define the empirical research program rather than sitting outside it. The strongest present claim is consequently narrower than a claim that CAPTUREDMIRAGE measures mastery: > **A chess-performance observation is interpretable only relative to the task, scoring rule, modeled dimension, support condition, cue state, exposure history, modality, opponent environment, occasion, and measurement process that produced it. A training system can make stronger later claims only if those distinctions survive collection and remain available for independent examination.** A system that preserves those distinctions can still reach the wrong conclusion. A system that discards them can make the stronger conclusion impossible to evaluate. The governing restriction follows: > **Do not let evidence obtained under one set of supports settle a competence claim whose truth depends on performance after those supports have been removed or changed.** The prospective test begins when that restriction is applied to CAPTUREDMIRAGE itself. --- # REFERENCES Barnett, S. M., & Ceci, S. J. (2002). When and where do we apply what we learn? A taxonomy for far transfer. *Psychological Bulletin, 128*(4), 612–637. https://doi.org/10.1037/0033-2909.128.4.612 Bilalić, M., McLeod, P., & Gobet, F. (2008a). Inflexibility of experts—Reality or myth? Quantifying the Einstellung effect in chess masters. *Cognitive Psychology, 56*(2), 73–102. https://doi.org/10.1016/j.cogpsych.2007.02.001 Bilalić, M., McLeod, P., & Gobet, F. (2008b). Why good thoughts block better ones: The mechanism of the pernicious Einstellung effect. *Cognition, 108*(3), 652–661. https://doi.org/10.1016/j.cognition.2008.05.005 Charness, N., Tuffiash, M., Krampe, R., Reingold, E., & Vasyukova, E. (2005). The role of deliberate practice in chess expertise. *Applied Cognitive Psychology, 19*(2), 151–165. https://doi.org/10.1002/acp.1106 Chase, W. G., & Simon, H. A. (1973). Perception in chess. *Cognitive Psychology, 4*(1), 55–81. https://doi.org/10.1016/0010-0285(73)90004-2 Corbett, A. T., & Anderson, J. R. (1995). Knowledge tracing: Modeling the acquisition of procedural knowledge. *User Modeling and User-Adapted Interaction, 4*(4), 253–278. https://doi.org/10.1007/BF01099821 de Groot, A. D. (1978). *Thought and Choice in Chess* (2nd ed.). De Gruyter Mouton. https://doi.org/10.1515/9783110800647 Ericsson, K. A., & Simon, H. A. (1993). *Protocol Analysis: Verbal Reports as Data* (Rev. ed.). MIT Press. https://doi.org/10.7551/mitpress/5657.001.0001 Ericsson, K. A., Krampe, R. T., & Tesch-Römer, C. (1993). The role of deliberate practice in the acquisition of expert performance. *Psychological Review, 100*(3), 363–406. https://doi.org/10.1037/0033-295X.100.3.363 Gobet, F., & Simon, H. A. (1996a). Recall of random and distorted chess positions: Implications for the theory of expertise. *Memory & Cognition, 24*(4), 493–503. https://doi.org/10.3758/BF03200937 Gobet, F., & Simon, H. A. (1996b). Templates in chess memory: A mechanism for recalling several boards. *Cognitive Psychology, 31*(1), 1–40. https://doi.org/10.1006/cogp.1996.0011 Gobet, F., & Simon, H. A. (1996c). The roles of recognition processes and look-ahead search in time-constrained expert problem solving: Evidence from grand-master-level chess. *Psychological Science, 7*(1), 52–55. https://doi.org/10.1111/j.1467-9280.1996.tb00666.x Karpicke, J. D., & Roediger, H. L. III. (2007). Repeated retrieval during learning is the key to long-term retention. *Journal of Memory and Language, 57*(2), 151–162. https://doi.org/10.1016/j.jml.2006.09.004 Krasny-Pacini, A., & Evans, J. (2018). Single-case experimental designs to assess intervention effectiveness in rehabilitation: A practical guide. *Annals of Physical and Rehabilitation Medicine, 61*(3), 164–179. https://doi.org/10.1016/j.rehab.2017.12.002 Macnamara, B. N., Hambrick, D. Z., & Oswald, F. L. (2014). Deliberate practice and performance in music, games, sports, education, and professions: A meta-analysis. *Psychological Science, 25*(8), 1608–1618. https://doi.org/10.1177/0956797614535810 Morris, C. D., Bransford, J. D., & Franks, J. J. (1977). Levels of processing versus transfer appropriate processing. *Journal of Verbal Learning and Verbal Behavior, 16*(5), 519–533. https://doi.org/10.1016/S0022-5371(77)80016-9 Office of Method. (2026a). *CAPTUREDMIRAGE 10.1.0 — Tournament Freeze: Release identity and build record* (O·M/CM-10.1/RI). Retained non-public primary artifact. Principal files: `RELEASE_IDENTITY.json`, `BUILD_REPRODUCTION_RESULTS.json`, `BUILD_PROVENANCE.md`. Office of Method. (2026b). *CAPTUREDMIRAGE 10.1.0 — Release assurance record* (O·M/CM-10.1/QA). Retained non-public primary artifact. Principal files: `CURRENT_RELEASE_CERTIFICATION.json`, `QA_RESULTS.txt`, and the retained release-assurance report. Office of Method. (2026c). *CAPTUREDMIRAGE 10.1.0 — Frozen curriculum and theory-provenance record* (O·M/CM-10.1/CURR). Retained non-public primary artifact. Principal files: `CURRICULUM_MANIFEST.md`, `THEORY_PROVENANCE.md`, `src/content/v10/manifest.json`. Office of Method. (2026d). *CAPTUREDMIRAGE 10.1.0 — Core evidence and Player Model algorithm* (O·M/CM-10.1/CORE). Retained non-public primary artifact. Principal file: `src/v10-core.js`. SHA-256: `8ca2a241f51fec6a8082b3340109e36a8f00b7d6dbfa9c6a36a57f7b85cbd4e9`. OpenTimestamps Project. (2026). *OpenTimestamps: A timestamping proof standard*. https://opentimestamps.org/ Rundgren, A., Jordan, B., & Erdtman, S. (2020). JSON Canonicalization Scheme (JCS). *RFC 8785*. RFC Editor. https://www.rfc-editor.org/rfc/rfc8785.html Roediger, H. L. III, & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. *Psychological Science, 17*(3), 249–255. https://doi.org/10.1111/j.1467-9280.2006.01693.x Shavelson, R. J., Baxter, G. P., & Gao, X. (1993). Sampling variability of performance assessments. *Journal of Educational Measurement, 30*(3), 215–232. https://doi.org/10.1111/j.1745-3984.1993.tb00424.x Shavelson, R. J., & Webb, N. M. (1991). *Generalizability Theory: A Primer*. Sage. Southwick, D. A., Harwell, K. W., Wright, G., Olsen, J. A., & Ogles, B. M. (2026). Not all practice is created equal: Longitudinal evidence from over 40,000 chess players. *Psychological Science, 37*(7), 514–523. https://doi.org/10.1177/09567976261452568 Stockfish Developers. (2026). *Stockfish 18 source code and UCI implementation*. https://github.com/official-stockfish/Stockfish Wang, C., Quan, Y., & Arthur, D. (2026). Review of cognitive diagnostic models (CDMs): Recent methodological advancements for addressing practical challenges. *British Journal of Mathematical and Statistical Psychology*. Advance online publication. https://doi.org/10.1111/bmsp.70066 --- # APPENDIX A — OPERATIONAL MATHEMATICS AND UNEQUAL SHRINKAGE ## A.1 Frozen context coefficients $$ q_c= \begin{cases} 0.42 & \text{guided},\\ 0.62 & \text{curriculum},\\ 0.76 & \text{mixed},\\ 0.88 & \text{OTB},\\ 0.92 & \text{sparring},\\ 1.00 & \text{serious game}. \end{cases} $$ ## A.2 Frozen quality coefficients | Grade | Coefficient | | ----- | ----------: | | A | 1.00 | | B | 0.84 | | C | 0.66 | | D | 0.46 | | E | 0.25 | ## A.3 Prompting, reliability, recency, and weight $$ p_i= \begin{cases} 0.78 & \text{prompted},\\ 1.00 & \text{otherwise}. \end{cases} $$ $$ r_i\in[0.1,1]. $$ Default half-life: $$ h_i=90\text{ days} $$ unless configured otherwise, subject to $$ h_i\ge21. $$ Optional event weight enters as $$ \max(0.25,\omega_i), $$ and the final implemented event-weight floor is $$ w_i\ge0.03. $$ ## A.4 Dimension-specific accumulator $$ \mathcal{I}_d = \{i:d_i=d,\;S_i\text{ defined}\}. $$ $$ \alpha_d = 2+ \sum_{i\in\mathcal{I}_d} w_iS_i. $$ $$ \beta_d = 2+ \sum_{i\in\mathcal{I}_d} w_i(1-S_i). $$ $$ \mu_d = \frac{\alpha_d} {\alpha_d+\beta_d}. $$ ## A.5 Weighted observed score $$ W = \sum_iw_i. $$ $$ p = \frac{\sum_iw_iS_i}{W}. $$ For symmetric prior \((a,a)\), $$ \mu_a(W;p) = \frac{a+pW}{2a+W} = p+\frac{a(1-2p)}{2a+W}. $$ ## A.6 Exact operational-gap decomposition For contexts \(G\) and \(T\), $$ G^{\mathrm{op}} = (p_G-p_T) + \frac{a(1-2p_G)}{2a+W_G} - \frac{a(1-2p_T)}{2a+W_T}. $$ If $$ p_G=p_T=p, $$ then $$ G^{\mathrm{op}} = \frac{ a(W_G-W_T)(2p-1) }{ (W_G+2a)(W_T+2a) }. $$ This demonstrates directly that the operational gap can contain shrinkage distortion even when no observed-score discrepancy exists. --- # APPENDIX B — RESEARCH EVIDENCE DATA DICTIONARY ### Table B1. Required research fields | Field | Meaning | | --------------------------------- | ------------------------------------------------------------------------------------------------------- | | Event ID | Immutable evidence identity | | Timestamp | Observation time | | Game ID | Game identity where applicable | | Round | Tournament round where applicable | | Ply | Exact decision ply | | Item/FEN ID | Exact task or position | | Source ID | Originating corpus/game/source | | Modeled dimension | Dimension credited/debited by operational event | | Raw outcome \(R_i\) | Native task outcome | | Process content reference \(P_i\) | Candidate/branch/reconstruction/process record where available | | Operational score \(S_i\) | Bounded dimension-specific operational evidence score where defined | | Outcome family | Binary / ordinal / bounded / continuous / categorical / other | | Outcome-rubric version | Exact scoring definition | | Skill taxonomy version | Canonical dimension ontology | | Mapping-codebook version | Pre-event task-demand mapping procedure | | Event item-mapping version | Post-event generated Q-row artifact where applicable | | Mapping confidence | Confidence in dimension relevance | | Mapping coder | Mapper identity | | Legacy operational context | Frozen compound context label | | Support facet | Instructional support state | | Cue facet | Thematic or other cue availability | | Exposure facet | First / repeated / consumed | | Modality | Digital / physical / other | | Opponent facet | None / training / human competition as applicable | | Competition facet | Practice / sparring / serious rated or analogous | | Quality grade | Frozen A–E operational grade | | Prompt state | Whether evidence was prompted | | Reliability | Implementation-defined reliability value | | Event modifier | Optional implementation-defined weight and reason | | Measurement mode | Concurrent / interaction-derived / pre-analysis retrospective / delayed retrospective / later inference | | Recall completeness | Complete-enough / partial / fragmentary / unknown | | Reconstruction delay | Minutes from game completion | | Reconstruction duration | Duration of primary reconstruction | | Next-round interval | Time remaining before next scheduled round | | Contamination state | Pre-analysis / contaminated / secondary | | Semantic state | Supported / quarantined / historical-only | | Exposure state | Active / sealed / consumed | | Benchmark epoch | Applicable assessment epoch | | Repair linkage | Repair case where applicable | | Corpus membership | Diagnostic A / Evaluation B / neither | | Sample universe | Objective all-moves / Corpus B / process-eligible subset | | Field Criterion Map version | Governing field opportunity/failure rubric | | Process coder | Research process coder | | Process codebook version | Governing process taxonomy | | Coding confidence | High / moderate / low | | Coding status | Primary / secondary / indeterminate | | Process-data exclusion reason | Why primary process inference is unavailable | | Engine manifest | Exact primary engine-analysis identity | | Engine-reference stability | Stable / escalated-stable / engine-indeterminate | | Missing-data reason | Reason required datum is absent | --- # APPENDIX C — CLAIM-STATUS REGISTER | Claim | Status | | ---------------------------------------------------------------- | ------------------------------------------------------------ | | 10.1.0 executable identity | Exact SHA-256 recorded | | Deterministic build repeatability | Two frozen-procedure builds produced identical shipping hash | | Semantic-content identity | Exact manifest SHA-256 recorded | | Active/sealed software separation | Internal contract/regression evidence | | Failed sealed retest cannot certify repair | Internal regression evidence | | Physical-board authority/reveal downgrade | Internal regression evidence | | 180-day state-transition simulation | 18,000 synthetic evidence events completed under release QA | | Context coefficients | Implemented operational policy | | Quality coefficients | Implemented operational policy | | Prompt multiplier 0.78 | Implemented operational policy | | Reliability clamp | Implemented operational policy | | Event-weight floors | Implemented operational policy | | Weighted beta-style accumulator | Implemented operational policy | | Operational prior sensitivity | Mathematically demonstrated | | Operational context coefficients psychologically calibrated | Not established | | Bounded operational encoding constitutes interval measurement | Explicitly rejected | | Item-to-skill mapping valid | Not yet validated | | Operational Transfer Gap | Implemented heuristic | | Operational Transfer Gap as transfer estimator | Inadmissible; unequal-shrinkage result demonstrated | | Operational warning prospective field criterion profile | Prospective | | Research-Plane context discrepancy | Exploratory/descriptive at Field Event A | | Mastery-stage predictive validity | Not established | | Field Criterion Map | Protocol-defined for Field Event A | | Corpus B process data missing at random | Not assumed | | Process-code inter-rater reliability | Not yet established | | Engine reference at 10M nodes always stable | Explicitly not assumed | | Physical-board modality effect | Prospective hypothesis | | Sealed-repair passage predicts lower recurrence | Prospective | | Adaptive scheduling improves training efficiency | Not established | | Tournament-performance effect | Not established | | Rating change attributable to CAPTUREDMIRAGE | Not established | | Global causal efficacy | Not established | | Public event identity required for pre-field scientific validity | Rejected; exact identity retained in embargoed lock | --- # APPENDIX D — FIELD EVENT A PROSPECTIVE PROTOCOL ## D.1 Public event facts **Designation:** Field Event A **Format:** Six-round in-person classical Swiss **Board:** Physical **Time control:** Classical with increment; exact control embargoed in the Pre-Event Lock **Exact identity, venue, dates, schedule, section, administrative identifiers:** Embargoed Pre-Event Lock ## D.2 In-game conduct No research notation during play. Only tournament-permitted chess notation is created. No engine, device, database, external analysis, candidate coding, confidence coding, or research annotation is used during the game. ## D.3 Primary post-game record Before analytical contamination: 1. preserve the legal game record; 2. record game-end and reconstruction-start time; 3. perform rapid chronological memory capture; 4. record candidates/resources/evaluations only where genuinely remembered; 5. record recall completeness; 6. for any endpoint intended for PRIMARY endpoint-evaluation coding, record `favorable`, `approximately balanced`, or `unfavorable` before analysis; 7. permit explicit “not remembered” entries; 8. stop at the applicable burden cap; 9. record reconstruction end; 10. freeze and timestamp the record. Burden caps: * ≥60 minutes until next round: 15 minutes; * 30–59 minutes: 8 minutes; * <30 minutes: no required primary process reconstruction. ## D.4 Corpus A Outcome-enriched diagnostic corpus. May use engine and post hoc chess analysis. Never used as event-wide process-prevalence estimator. ## D.5 Corpus B Eligibility: * participant to move; * played game; * fullmove ≥8; * at least two legal moves; * legal reconstructable position. Temporal strata: * A: fullmoves 8–18; * B: fullmoves 19–35; * C: fullmoves 36+. Quota: * maximum two positions per game/stratum; * no cross-stratum backfill. Ranking payload: `OM04-WP001-FIELD-A-B2|round=|ply=

` Encoding: UTF-8, no trailing newline. Hash: SHA-256. Selection: two lexicographically smallest digests per game/stratum. ## D.6 Mapping Pre-event frozen: * taxonomy; * mapping codebook; * field-assessability rules; * confidence rubric. Post-event mapping packet contains only: * sampled-position identity; * FEN; * side to move; * complete legal-move list. It excludes the participant move, warning state, process reconstruction, and engine evaluation of the played move. Q rows are generated and hashed before those outcome fields are revealed. When the author maps, this is sequential information masking, not blinding. ## D.7 PRIMARY field frame PRIMARY Field Event A dimensions: * candidate-generation; * opponent-resistance; * endpoint-evaluation; * comparison. `branch-control` and `move-order` are CONDITIONAL rather than PRIMARY in Version 1.0. The normative SUCCESS/FAILURE/INDETERMINATE rules are Table 5 and Appendix H. ## D.8 Opportunity accounting $$ N_d^B = \text{all Corpus B field opportunities for }d. $$ $$ I_d^B = \text{indeterminate opportunities}. $$ $$ C_d^B = N_d^B-I_d^B. $$ $$ F_d^B = \text{failures among classifiable opportunities}. $$ $$ R_d^B = F_d^B/C_d^B \quad (C_d^B>0). $$ Low-confidence putative process classifications are INDETERMINATE and remain inside \(N_d^B\). All quantities are reported. ## D.9 Pre-event warning snapshot Snapshot is frozen only after final substantive pre-event CAPTUREDMIRAGE training ends. Exact absolute timestamp is recorded in the embargoed lock. No scored CAPTUREDMIRAGE training or targeted remediation occurs afterward before the first round. Any violation creates `INTERVENED-BEFORE-CRITERION` status for affected dimensions. ## D.10 Engine protocol Stockfish 18. Discovery: * Threads = 1; * Hash = 512 MB; * MultiPV = 5; * `UCI_ShowWDL = true`; * `UCI_Chess960 = false`; * 10,000,000 nodes. Adjudication: * each Stage-A discovery move and each explicitly recalled human candidate receives an independent single-root `searchmoves` search at 10,000,000 nodes; * all material comparison values are taken from these symmetric single-root searches; * the same symmetry applies at 50,000,000-node escalation. Before each search: `ucinewgame` → `isready` → `readyok` → `position fen` → `go ...` The lock records the exact executable, hashes, build flavor, OS/version, CPU, UCI option dump, startup output, all loaded NNUE filenames/hashes, Syzygy identities, and parser identity. ## D.11 Ordinary WDL expected score $$ E_{\mathrm{WDL}} = \frac{W+\tfrac12D}{1000}. $$ $$ \Delta E = E_{\mathrm{best}}-E_{\mathrm{candidate}}. $$ Ordinary candidate acceptable only when $$ \Delta \mathrm{CP}\le50 $$ and $$ \Delta E\le0.05. $$ ## D.12 Tablebase and forced-mate rules **Tablebase-eligible:** primary candidate acceptability is preservation of the exact root-side WDL class under the locked 50-move-rule convention. DTZ is descriptive for the PRIMARY candidate-generation endpoint. **Forced mate:** a candidate is PRIMARY-acceptable when the locked symmetric search also preserves forced mate for the root side. Mate distance is descriptive. Unstable forced-mate status yields `ENGINE-REFERENCE INDETERMINATE`. ## D.13 Stability escalation Escalate where: $$ 35\le\Delta\mathrm{CP}\le65 $$ or $$ 0.035\le\Delta E\le0.065. $$ If symmetric 10M and 50M classifications disagree: `ENGINE-REFERENCE INDETERMINATE`. ## D.14 Warning–criterion profile For each PRIMARY dimension, report: $$ (A_d,N_d^B,I_d^B,C_d^B,F_d^B,R_d^B). $$ No common cross-dimensional concordance coefficient, discrimination statistic, accuracy statistic, or success threshold is defined. If all PRIMARY dimensions share the same alert state, no alert-versus-non-alert comparison exists. ## D.15 Coding If a qualified independent coder is secured before the lock, every primary classifiable case submitted to the main process analysis is independently coded. Any technically unavailable case is identified individually and excluded from the complete-case inter-rater statistic. Otherwise: * no inter-rater reliability is claimed; * delayed self-recoding is labeled intra-rater consistency only. ## D.16 Reporting No tournament-level significance testing. Report objective and process evidence with denominators, missingness, mapping, engine status, warning status, intervention boundaries, and deviations. --- # APPENDIX E — RELEASE IDENTITY AND BUILD RECORD **Product:** CAPTUREDMIRAGE **Release:** 10.1.0 **Release name:** TOURNAMENT FREEZE **Release date:** 2026-09-02 **Semantic-content epoch:** 10.0 **Profile schema:** 18 **v10 state schema:** 1 **Shipping executable:** `CAPTUREDMIRAGE.exe` **Platform:** Windows amd64 / PE32+ **Size:** 113,703,424 bytes **Executable SHA-256** `ea450fe794f6c6f963747ee5b6221838dfd617c4bb45173754bdf719828c44c8` **Release-package SHA-256** `89c7c6977fec84d966b5fdcfc00b32e6fa94647a6d92221868dc5a8151d69d0a` **Frozen semantic-content manifest SHA-256** `b4eeb1cecd9011e5111de14110c4348b2de5a9440ff6ce0d818512c91f21a69f` **Historical source manifest SHA-256** `799f3396a387fffa29cc764c6152a250284ef5a7602ce5cbb2fea0006249c9a6` **Core algorithm `src/v10-core.js` SHA-256** `8ca2a241f51fec6a8082b3340109e36a8f00b7d6dbfa9c6a36a57f7b85cbd4e9` **Release identity record SHA-256** `e76520ed9c4471e6d1f1c74aeb90a3b58841a32fd278606df468f08cd895fde9` **Curriculum manifest record SHA-256** `f158e2d15ad8cfdc0e6c3400f774d9e9765e865b7c76baf59627eebad63bc5a7` **Release assurance report SHA-256** `5b2570602fd9582cd10bcc3e120f640c99f6bab865cadd869bb8a1b16564dc74` **Current release certification SHA-256** `a89e2a5a3e90d787cb86a09ad16b7a701d1cc5bc7f412f2eae75616ad75a8f51` **QA results SHA-256** `999cb21aa74a54f44873fd5167350c898358453847afb8dd1722b289ada2ac2a` **Embedded Stockfish 18 archive SHA-256** `40cc975817e7eee270b03f354810d20956df565420d320f6dd37d454dc81a139` **Frozen compatibility identities** * v10.0.0: `610203642c294503de2db322ed09d990dec60a0785f6217a83407fee8df3e743` * v9.0.0: `14c667d6a721d52802223d28c4d6e5fbcab3266af8ba5e43e74a349df3d7378b` * v8.0.1: `99219185e7e8bbbe2010539627e381b52d452b5171b435aa299591df9a791f9d` **Build statement:** The Windows host was built from the frozen v10.1 source with deterministic-build controls including `CGO_ENABLED=0`, Windows amd64 targeting, `-trimpath`, disabled VCS stamping, and empty build ID. Two separate builds under the frozen procedure produced the same shipping executable hash. **Release-assurance boundary:** All 28 enumerated current-state executable QA programs/contracts passed. Managed-browser navigation remained outside the executable gate under environment policy. Authenticode is absent. Maia is not bundled. The internal release identity contains campaign-specific administrative metadata that is intentionally not reproduced in the public manuscript. Its cryptographic identity remains published above. --- # APPENDIX F — PRIMARY FIELD PROCESS CODEBOOK The operational failure taxonomy remains larger than the research coding set. The process codes below localize supported divergences. PRIMARY field outcomes for the four PRIMARY dimensions are governed by Table 5; a process code alone does not automatically create a quantitative failure. ## F1. Representation divergence **Definition:** The explicit process reconstruction supports that the participant reasoned from an incorrect board state or legal relation that changes the legality, tactical status, or objective reference of the analyzed decision. **Required evidence:** Specific misplacement, stale relation, omitted piece, incorrect side to move, hallucinated blocker, missed opened/closed line, or equivalent representational discrepancy. **Exclusion:** Mere failure to notice a tactic without evidence that the represented board state itself was wrong. **Reference:** Legal reconstructed board state. **Primary warning-map status:** Diagnostic process code; not itself a PRIMARY Player Model field dimension in Field Event A. ## F2. Candidate omission **Definition:** No explicitly recorded candidate belongs to the frozen primary acceptable-candidate set in a process record sufficiently complete to support exclusion. **Required evidence:** `recall_completeness = complete-enough`, legal candidate identities, and stable candidate-complete symmetric reference. **Exclusion:** Acceptable candidate considered but rejected later; partial/fragmentary/unknown process record. **Mapped PRIMARY dimension:** `candidate-generation`. ## F3. Opponent-resistance omission **Definition:** In a branch record sufficiently complete to support negative inference, no explicitly recalled opponent reply belongs to the frozen principal-resistance set at the relevant opponent node. **Required evidence:** Complete-enough branch reconstruction, legal reply identities, and stable symmetric opponent-node reference. **Exclusion:** Principal resource considered but later mis-evaluated; incomplete retrospective record. **Mapped PRIMARY dimension:** `opponent-resistance`. ## F4. Branch-control or termination divergence **Definition:** The record supports premature branch termination before a consequential continuation was considered. **Status in Field Event A:** CONDITIONAL. No general PRIMARY quantitative rule is frozen in Version 1.0 because “sufficient continuation” cannot be reduced here to a non-arbitrary universal ply count or engine threshold. ## F5. Endpoint-evaluation divergence **Definition:** A legally reconstructed branch endpoint has an explicit pre-analysis root-side label and that label is incompatible with the frozen objective three-way band under Table 5. **Mapped PRIMARY dimension:** `endpoint-evaluation`. **Primary FAILURE:** opposite-polarity disagreement only: favorable vs objective unfavorable, or unfavorable vs objective favorable. **Adjacent-band disagreement:** INDETERMINATE for PRIMARY quantitative reporting. ## F6. Comparison/selection divergence **Definition:** Among at least two explicitly analyzed candidates with recorded comparative preference or selection, the preferred candidate is materially dominated by another explicitly analyzed candidate under both frozen objective thresholds. **Mapped PRIMARY dimension:** `comparison`. **Primary SUCCESS:** the preferred/selected candidate belongs to the frozen acceptable-candidate set relative to the best explicitly analyzed candidate. **Primary FAILURE:** the preferred/selected candidate falls outside that acceptable set because another explicitly analyzed candidate is better by more than 50 cp **and** more than 0.05 root-side expected score under symmetric adjudication. **Tie/equivalence region:** acceptable candidates inside the frozen equivalence thresholds count as SUCCESS; incomplete ranking or unstable reference is INDETERMINATE. ## F7. Move-order divergence **Definition:** The participant identifies a sequence-sensitive operation but chooses an order that appears inferior to another implementation of the same intended operation. **Status in Field Event A:** CONDITIONAL. No PRIMARY quantitative rule is frozen in Version 1.0 because identifying “the same intended operation” would otherwise require post-event strategic discretion. ## F8. Final recheck/verification divergence **Definition:** The chosen move or line has otherwise survived the recorded decision process, but a final legality, forcing-resource, or tactical verification that the process expected was omitted or ineffective. **Status in Field Event A:** CONDITIONAL; not part of the PRIMARY comparison frame. ## F9. Multiple supported divergences — ordering indeterminate **Definition:** Two or more process divergences are supported, but the evidence does not establish which occurred first. No artificial earliest-stage attribution is imposed. ## F10. Process indeterminate **Definition:** Available evidence is insufficient to locate a process divergence defensibly. This is a legitimate primary outcome. ### Absence-of-mention rule Silence in a retrospective account is not sufficient evidence that a candidate, resource, or thought was absent during play. Negative omission codes require a record explicitly classified as complete enough to support exclusion. ### Confidence Every non-indeterminate process code receives a confidence judgment concerning **evidentiary support for the code**, not certainty about the participant's unobservable cognition. **High:** explicit pre-analysis process reconstruction plus strong, stable chess-reference support and no comparably plausible competing code. **Moderate:** process reconstruction and chess evidence support the code, but a plausible alternative remains. **Low:** tentative attribution. A low-confidence putative classification remains an opportunity in \(N_d^B\) when opportunity criteria are met, but it is coded INDETERMINATE and contributes to \(I_d^B\) rather than being removed from the denominator. --- # APPENDIX G — PRE-EVENT LOCK, TIMESTAMPING, AND MATERIALS AVAILABILITY ## G.1 Commitment architecture The Pre-Event Lock is defined by `PRE_EVENT_LOCK_MANIFEST.json` rather than by an archive container. For every included artifact the manifest records: * normalized relative path; * byte size; * SHA-256; * semantic role. The manifest is serialized using the JSON Canonicalization Scheme specified in RFC 8785. The exact RFC-8785 canonical UTF-8 bytes are the authoritative manifest bytes. Define $$ H_{\mathrm{lock}} = SHA256(\text{exact RFC-8785 canonical manifest bytes}). $$ The lock records the exact absolute ISO 8601 timestamp of the final prediction snapshot and the event-administrative details withheld from the public manuscript. ## G.2 External timestamp OpenTimestamps is run directly against the exact canonical `PRE_EVENT_LOCK_MANIFEST.json` file. The `.ots` proof therefore attests that file rather than a separately created textual representation of its SHA-256 digest (OpenTimestamps Project, 2026). The lock records: * OpenTimestamps client/tool identity and version; * initial stamping time; * initial proof status; * later proof-upgrade status where applicable; * later verification status where applicable. The separately reported \(H_{\mathrm{lock}}\) is the human-readable SHA-256 identity of the same canonical file bytes. The commitment mechanism permits the identifying file contents to remain embargoed while preserving a verifiable pre-event existence claim. ## G.3 Version 1.0 publication package The Version 1.0 publication package consists of: * canonical paper HTML; * canonical downloadable TXT source; * canonical downloadable PDF; * `OM-04-WP-001-v1.0-MANIFEST.json`, recording the SHA-256 identities of those publication artifacts; * version-history statement; * public CAPTUREDMIRAGE release identities used in the paper; * public description of the Pre-Event Lock procedure. The manuscript does not embed its own SHA-256 because doing so would alter the bytes being identified. Publication-artifact identities are authoritative in the release manifest. ## G.4 Hash-identified but retained non-public The following can remain non-public while their identities are disclosed: * CAPTUREDMIRAGE 10.1.0 executable; * source tree; * internal QA bundle; * detailed release-assurance artifacts; * proprietary or operational research data not necessary for reproducing the paper's argument. A cryptographic hash proves identity. It does not create independent inspectability. The paper therefore does not use *publicly reproducible* as a synonym for *hash-identified*. ## G.5 Embargoed before the field event The following remain embargoed because public disclosure would create unnecessary identity correlation: * exact Field Event A name; * venue; * exact dates; * exact schedule and time control; * section; * registration evidence; * participant federation/rating identifiers; * exact event-administrative rule record; * full Pre-Event Lock contents. The root commitment establishes pre-event fixation without requiring immediate disclosure of those fields. ## G.6 Deviations Every post-lock deviation records: * timestamp; * reason; * affected rule/artifact; * knowledge state at the time of deviation; * effect on primary versus secondary analysis. The original lock is immutable. A deviation does not rewrite the preregistration. --- # APPENDIX H — FIELD CRITERION MAP FOR ALL 43 PLAYER MODEL DIMENSIONS The v10.1 skill tree contains 43 modeled dimensions across strategy, calculation, endgame, opening, transfer, evaluation, and style. Field Event A does **not** pretend that all 43 can be measured validly from six tournament games. The following classification is frozen before the event. The PRIMARY set is exactly `candidate-generation`, `opponent-resistance`, `endpoint-evaluation`, and `comparison`; all other dimensions are CONDITIONAL or NOT PRIMARY FIELD-ASSESSABLE for Version 1.0. ### Status codes **PRIMARY** — eligible for the primary prospective warning–criterion profile under an explicit SUCCESS/FAILURE/INDETERMINATE rule. **CONDITIONAL** — can yield descriptive or secondary field evidence only when the listed evidentiary condition is satisfied. **NOT PRIMARY FIELD-ASSESSABLE** — remains an operational dimension but no defensible direct failure criterion is assigned for this event. ### H.1 Strategy | Dimension | Status | Field opportunity / criterion | | -------------------------- | ----------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | structure-recognition | CONDITIONAL | Requires blind task-demand mapping showing a dominant structural classification plus an explicit pre-analysis participant representation. Explicit materially incorrect structure can support failure; silence cannot. | | attention-prioritization | CONDITIONAL | Requires explicit record of what was treated as positionally critical and a frozen expert/reference rubric. No omission inferred from silence. | | trajectory-selection | CONDITIONAL | Requires explicit strategic trajectory statement and later stable chess reference. Descriptive only in Field Event A. | | opponent-intent | CONDITIONAL | Requires explicit opponent-plan representation in a sufficiently complete process account and a reference showing a materially relevant opposing plan. | | transformation-forecasting | CONDITIONAL | Requires explicit forecast of structural/material transformation and objective post-position reference. | | break-timing | CONDITIONAL | Requires a mapped break opportunity and explicit decision rationale; no universal field-failure threshold is imposed. | | exchange-consequence | CONDITIONAL | Requires explicit exchange assessment and stable consequence reference. | | whole-board-awareness | CONDITIONAL | Only explicit representational/attention failures can be coded; absence of mention does not establish absence of awareness. | ### H.2 Calculation | Dimension | Status | Field opportunity / criterion | | -------------------- | ----------- | ------------------------------------------------------------------------------------------------------------------------------------- | | candidate-generation | **PRIMARY** | Corpus B position mapped as candidate-generation relevant; process record complete enough for candidate-set inference. Failure = F2. | | candidate-ordering | CONDITIONAL | Requires reliable recalled discovery/analysis order; retrospective order uncertainty prevents primary use. | | opponent-resistance | **PRIMARY** | Opportunity/scoring follows Table 5: complete-enough branch record plus symmetric opponent-node reference; SUCCESS/FAILURE/INDETERMINATE determined by principal-resistance-set membership. | | branch-control | CONDITIONAL | Process divergence can be described under F4, but no general non-arbitrary PRIMARY quantitative endpoint is frozen for Version 1.0. | | endpoint-evaluation | **PRIMARY** | Opportunity/scoring follows Table 5: legal endpoint plus pre-analysis favorable/balanced/unfavorable label and stable objective three-way reference. | | comparison | **PRIMARY** | Opportunity/scoring follows Table 5: at least two explicitly analyzed candidates with recorded preference/selection and stable symmetric references. | | move-order | CONDITIONAL | Process divergence can be described under F7, but “same intended operation” is not given a universal PRIMARY quantitative rule in Version 1.0. | | decision-sufficiency | CONDITIONAL | Can be described where the participant explicitly reports stopping criteria and the reference shows unresolved material; not primary. | | time-allocation | CONDITIONAL | Requires reliable per-move timing and a frozen decision-class time rubric. No timing data means no field score. | ### H.3 Endgame | Dimension | Status | Field opportunity / criterion | | ----------------------- | ----------- | -------------------------------------------------------------------------------------------------------------------------------------- | | exact-anchor-recall | CONDITIONAL | Requires position matching a frozen exact/theoretical anchor family and explicit recall evidence. | | mechanism-transfer | CONDITIONAL | Requires mapped endgame mechanism in an unrelated field position and explicit process evidence. | | critical-move-execution | CONDITIONAL | Requires exact or stable engine/tablebase narrow-move reference. Objective execution can be described; process cause remains separate. | | endgame-transition | CONDITIONAL | Requires mapped transition opportunity and explicit transition judgment. | | king-access | CONDITIONAL | Requires endgame position where king-access relation is materially relevant under the frozen endgame codebook. | | rook-activity | CONDITIONAL | Requires rook-endgame opportunity and explicit activity judgment or objective execution criterion. | | fortress-recognition | CONDITIONAL | Requires objectively supportable fortress/non-fortress status and explicit recognition evidence. | | conversion-technique | CONDITIONAL | Requires objectively winning endgame and sufficient sequence to evaluate conversion; six-game event may supply no opportunity. | | defensive-technique | CONDITIONAL | Requires defensible/drawable endgame with sufficiently stable exact/reference evidence. | ### H.4 Opening | Dimension | Status | Field opportunity / criterion | | -------------------- | ----------- | -------------------------------------------------------------------------------------------------------------------------- | | opening-recall | CONDITIONAL | Opening-specific analysis outside generic Corpus B; requires exact repertoire-decision mapping and explicit recall status. | | opponent-deviation | CONDITIONAL | Requires opponent departure from mapped repertoire and explicit recognition/response evidence. | | theory-exit | CONDITIONAL | Requires identifiable boundary between frozen repertoire knowledge and independent decision. | | opening-to-structure | CONDITIONAL | Requires explicit structural handoff after theory exit. | ### H.5 Transfer | Dimension | Status | Field opportunity / criterion | | --------------------- | ---------------------------- | ---------------------------------------------------------------------------------------------------------------------- | | mixed-transfer | NOT PRIMARY FIELD-ASSESSABLE | Operational evidence-context dimension rather than independent tournament process failure. | | cue-free-transfer | NOT PRIMARY FIELD-ASSESSABLE | Operational evidence-context dimension; Field Event A itself supplies cue-free serious-game evidence where applicable. | | otb-transfer | NOT PRIMARY FIELD-ASSESSABLE | Modality/context dimension, not a direct process-failure criterion in this event. | | sparring-transfer | NOT PRIMARY FIELD-ASSESSABLE | Training-context dimension, not a Field Event A failure criterion. | | serious-game-transfer | NOT PRIMARY FIELD-ASSESSABLE | Serious-game evidence state; should not be circularly scored from itself as a primary field failure. | ### H.6 Evaluation | Dimension | Status | Field opportunity / criterion | | ------------------------ | ----------- | ----------------------------------------------------------------------------------------------------------------------------------------- | | wdl-calibration | CONDITIONAL | Requires explicit pre-analysis WDL/evaluation judgment and stable objective reference. | | static-dynamic-character | CONDITIONAL | Requires explicit characterization and frozen reference rubric. | | confidence-calibration | CONDITIONAL | Requires contemporaneously or legitimately reconstructed confidence evidence; tournament protocol creates no in-game confidence notation. | ### H.7 Style | Dimension | Status | Field opportunity / criterion | | ---------------------- | ---------------------------- | -------------------------------------------------------------------------------------------- | | liberation-suppression | NOT PRIMARY FIELD-ASSESSABLE | Style-governor signal; may be described post hoc but not used as primary criterion endpoint. | | dynamic-conversion | NOT PRIMARY FIELD-ASSESSABLE | Style-governor signal rather than independent field process criterion. | | anti-passivity | NOT PRIMARY FIELD-ASSESSABLE | Style-governor signal; no primary field failure rule assigned. | | anti-dogma | NOT PRIMARY FIELD-ASSESSABLE | Normative style dimension; no primary tournament criterion. | | plan-freshness | NOT PRIMARY FIELD-ASSESSABLE | Normative style dimension; descriptive only unless a later dedicated protocol is created. | ### H.8 Governing rule Only **PRIMARY** dimensions enter the main prospective warning–criterion profile for Field Event A. Conditional dimensions can generate secondary evidence when their preregistered evidentiary conditions happen to be met. Dimensions marked **NOT PRIMARY FIELD-ASSESSABLE** remain legitimate operational dimensions in CAPTUREDMIRAGE. Their exclusion from the first field study reflects limits of observation, not a claim that they are unimportant. A pre-event operational warning falling entirely outside the PRIMARY field-assessable frame is reported as **not assessable under this event protocol**. It is not replaced after the tournament with a more convenient dimension.