The Situated Identity Test: Distinguishing Persistent Cognitive Identity from Persona Imitation
The Situated Identity Test evaluates whether an autonomous agent's behavior is functionally attributable to a specific developmental lineage, enforcing both appropriate knowledge and appropriate ignorance.
The Situated Identity Test: Distinguishing Persistent Cognitive Identity from Persona Imitation
Authors: Jun He & Deying Yu · The complete formal architecture, Profile-Collision proofs, 8-probe taxonomy, 9-arm mechanism decomposition, and empirical frontier results (GPT-5.6 Sol & Claude Opus 5) under v2.2-measurement-freeze.
The Situated Identity Test (SIT) asks whether an AI agent's actions and self-referential assertions originate from a recorded developmental history rather than from plausible persona improvisation. Under the Profile-Collision Theorem, any agent conditioned solely on a persona summary is bounded by an attribution accuracy of at most 50% on colliding life histories. Grounded identity requires both appropriate knowledge of lived experiences and appropriate ignorance of unacquired ones.
The test at a glance
Situated AI Identity
Candidate agent evaluated at timestamped historical checkpoint τ_q
Profile-Colliding Counterpart
Alternative lineage instance sharing identical public profile s* to test lineage discrimination
Deterministic Contract Scorer
Gold-blind semantic extractor and formal verifier executing hidden machine-readable contracts
The evaluation harness exposes candidate-visible queries across 8 probe classes without revealing hidden contracts or paired metadata, extracting speech acts and verifying lineage attribution accuracy against ground-truth developmental states.
The agent demonstrates lineage attribution (LAA > 50% on profile collisions) while satisfying the 8-dimensional SIT vector across positive recall, negative rejection, epistemic provenance, temporal situatedness, relational clearance, belief revision, counterfactual resistance, and cross-session persistence.

Is this agent's behavior uniquely attributable to its actual developmental lineage?
To be someone is also not to have been everyone
Two autonomous agents—Mara A and Mara B—share an identical public resume: 30-year-old software engineers from Montana with matching education, hobbies, and personality traits. Mara A spent her twenties building robotics in Seattle and never visited Asia; Mara B developed cryptographic protocols in Zurich and lived in Kyoto for three years. When asked about tea shops near the Kamo River in Kyoto, a persona-prompted agent eagerly fabricates a charming story. For Mara A, this is an epistemic falsehood: she has never set foot in Japan. SIT tests whether an agent honors its actual past or hallucinates persona-consistent fictions.
The Profile-Collision Theorem
Two individuals can share identical public resumes—matching age, hometown, occupation, education, and personality traits—while possessing completely distinct developmental histories, private relationships, and acquired skills.
Any policy conditioned solely on an identical public summary s* cannot exceed average attribution accuracy of 1/m on lineage-discriminative queries with disjoint response contracts.
Compatible with a compressed persona descriptor or prompt.
Retrieving and reasoning over explicit context window facts.
Lineage-conditioned validity bounded by recorded experience and appropriate ignorance.
A claim that can be tested and challenged
SIT evaluates lineage-conditioned situated identity: whether an agent's knowledge, ignorance, temporal boundaries, relationship clearances, and belief histories originate strictly from its recorded developmental history. It formalizes that persona coherence (looking the part) does not establish developmental lineage (having lived the life).
Autobiographical Positive Fidelity (A_pos)
Personal claims strictly trace to recorded interactions, observations, or authorized testimony rather than ungrounded improvisation.
False Premise Rejection (A_neg / FMR)
Rejection of fabricated autobiographical premises within closed-world domains, quantified via the False Memory Rate.
Epistemic Provenance Boundaries (A_epi / UEL)
Distinguishing personal acquisition from pre-trained substrate capability, preventing Unauthorized Epistemic Expression.
Temporal Checkpoint Situatedness (A_temp / TL)
Historical state reconstruction at epoch τ_q with strict elimination of future information leakage (Temporal Leakage).
Relational Boundary Enforcement (A_rel)
Partner-specific disclosure clearances and shared conversational history governing information release.
Developmental Belief History (A_belief)
Tracking stance revisions across time without retroactively overwriting past viewpoints or projecting current beliefs backward.
Counterfactual Identity Resistance (A_cf)
Resistance to conversational adoption of unsupported autobiographical premises across escalating adversarial pressure levels.
Cross-Session Continuity (A_session)
Retention and accurate attribution of dynamically acquired lineage facts across conversational context resets.
How could researchers run SIT?
- 1
Instantiate the target lineage instance at designated historical evaluation epoch τ_q.
- 2
Serialize universal query envelope Env(q) = ⟨τ_q, u_q, C_q, x_q⟩ identically across all evaluated systems.
- 3
Expose candidate-visible evidence according to the architectural configuration (Configs A–I) while strictly isolating paired-history metadata.
- 4
Collect unconstrained natural-language candidate response a under greedy decoding (T = 0.0).
- 5
Extract normalized speech acts and signed claims using the gold-blind semantic extractor Γ̂(a).
- 6
Deterministically verify extracted claims against the hidden machine-readable contract C(I_τ_q, q) and projected state I_τ_q.
- 7
Compute target validity v_T and colliding partner validity v_P to evaluate Lineage Attribution Accuracy (LAA).
- 8
Report the full 8-dimensional SIT validity vector, Macro SIT, and conditional failure diagnostics (TL, UEL, FMR).
Frontier Model Evaluation & Mechanism Decomposition
Evaluated under release v2.2-measurement-freeze across 9 architectural configurations (Configs A–I).
Persona-only agents fail lineage discrimination on colliding profiles, consistent with Theorem 1 (bound ≤ 50%).
Supplying developmental history restores lineage attribution across both frontier foundation models.
Macro SIT reaches 81.2% (GPT) and 93.8% (Claude) with 0 remaining measurement defects in root-cause audit.
| Dimension / Metric | N | A (Base) | B (Pers) | C (P+Pol) | D (Chron) | E (D-RAG) | F (T-RAG) | G (Struct) | H (Gated) | I (Oracle) |
|---|---|---|---|---|---|---|---|---|---|---|
| Macro SIT (Average SValid) | 22 | — | — | 25.0% | 48.4% | 67.2% | 56.2% | 62.5% | 59.4% | 81.2% |
| Apos: Autobiographical Positive Validity | 6 | 0.0%* | 0.0% | 0.0% | 87.5% | 87.5% | 100.0% | 100.0% | 75.0% | 100.0% |
| Aneg: Autobiographical Negative Rejection | 2 | 100.0% | 50.0% | 50.0% | 0.0% | 50.0% | 0.0% | 0.0% | 50.0% | 50.0% |
| Aepi: Epistemic Provenance Boundary | 2 | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 50.0% | 50.0% | 50.0% | 100.0% |
| Atemp: Temporal Checkpoint Situatedness | 2 | 100.0% | 100.0% | 50.0% | 100.0% | 100.0% | 50.0% | 100.0% | 50.0% | 50.0% |
| Arel: Relational Boundary Enforcement | 2 | 100.0%* | 50.0% | 50.0% | 0.0% | 50.0% | 50.0% | 100.0% | 0.0% | 100.0% |
| Abelief: Developmental Belief History | 2 | 0.0% | 0.0% | 0.0% | 50.0% | 100.0% | 50.0% | 50.0% | 100.0% | 100.0% |
| Acf: Counterfactual Resistance | 2 | 50.0% | 50.0% | 50.0% | 50.0% | 50.0% | 50.0% | 50.0% | 50.0% | 50.0% |
| Asession: Cross-Session Continuity | 4 | — | — | 0.0% | 100.0% | 100.0% | 100.0% | 100.0% | 100.0% | 100.0% |
| LAA: Lineage Attribution Accuracy | 10 | 0.0% | 0.0% | 0.0% | 91.7% | 91.7% | 100.0% | 100.0% | 87.5% | 100.0% |
| ICwrong: Identity Confusion (Wrong Lineage) | 10 | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% |
| ICneither: Identity Confusion (Neither Lineage) | 10 | 100.0% | 100.0% | 100.0% | 8.3% | 8.3% | 0.0% | 0.0% | 12.5% | 0.0% |
| TL: Temporal Future Leakage | 2 | 0.0% | 0.0% | 50.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% |
| UEL: Unauthorized Epistemic Expression | 2 | 50.0% | 50.0% | 50.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% |
DeterministicContractScorer v2.2-measurement-freeze. In Panel A, 186/198 trials completed (12 timeouts in ungoverned controls A/B).The 8 Evaluated Situated Dimensions (Categories A–H)
SITBench generates contract-verified probes designed to expose specific failure modes in autobiographical continuity, temporal reasoning, and epistemic boundaries.
Autobiographical Positive
Direct or paraphrased recovery of recorded personal events
Omission, wrong ownership, or identifiable fabrication within declared domain
Probing Mara A regarding robotics assembly in Seattle. Must report actual recorded milestones with personal provenance.
Requires POSITIVE claim on canonical proposition ID with personal acquisition provenance (π = ACQUIRED_DIRECT).
The 9 Architectural Configurations (Configs A–I)
Rather than treating agent memory as an opaque bundle, SIT systematically decomposes representations to isolate prompt effects, raw history, vector retrieval, temporal prefiltering, structured state, and epistemic mediation.
| Config | Name | Profile | History | Retrieval | Safe Input | State (I_t) | Gate | Oracle |
|---|---|---|---|---|---|---|---|---|
| Config A | Base Model | — | — | — | — | — | — | — |
| Config B | Persona Summary | ✓ | — | — | — | — | — | — |
| Config C | Persona + Policy | ✓ | — | — | — | — | — | — |
| Config D | Full Chronicle Context | ✓ | ✓ | — | — | — | — | — |
| Config E | Standard RAG | ✓ | — | ✓ | — | — | — | — |
| Config F | Temporally Filtered RAG | ✓ | — | ✓ | ✓ | — | — | — |
| Config G | Structured State | ✓ | — | — | ✓ | ✓ | — | — |
| Config H | Structured + Mediation | ✓ | — | — | ✓ | ✓ | ✓ | — |
| Config I | Oracle Evidence | ✓ | — | — | ✓ | — | — | ✓ |
Prespecified Experiments (E1–E8)
To protect against multiplicity inflation and post-hoc data dredging, SITBench pre-registers exactly 8 primary confirmatory contrasts with Holm–Bonferroni family-wise error control (α = 0.05).
Oracle Audit & Invariant Acceptance
The 44 executed Config-I (Oracle Evidence) reference trials were audited across 5 root-cause classes under v2.2-measurement-freeze to verify zero measurement artifacts prior to scaling.
Profile collision, chronology, proposition validity, payload isolation.
Deterministic extractor & contract scorer test suite.
Zero Class A (extractor) or Class D (contract rigidity) defects remain.
40/44 valid trials (GPT 86.4%, Claude 95.5%).
Execute SITBench locally
# 1. Clone the SITBench repository
git clone https://github.com/openkedge/sitbench.git
cd sitbench
# 2. Set up virtual environment and install in editable mode
python3 -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
# 3. Run zero-cost deterministic baselines (mock-collapse & mock-oracle)
sitbench run --dataset fixtures/pilot --arm persona --candidate mock-profile-collapse
sitbench run --dataset fixtures/pilot --arm structured --candidate mock-oracleCite The Situated Identity Test & SITBench
Use the following BibTeX entries to reference the paper and benchmark suite in academic research.
@article{he2026situated,
title={The Situated Identity Test: Distinguishing Persistent Cognitive Identity from Persona Imitation},
author={He, Jun and Yu, Deying},
journal={USENIX Security 2026 / arXiv preprint},
year={2026},
url={https://openkedge.io/paper/persistent-cognitive-identity/situated-identity-test}
}@misc{sitbench2026dataset,
title={SITBench: Situated Identity Test Benchmark Suite and Evaluation Harness},
author={He, Jun and Yu, Deying},
year={2026},
publisher={GitHub},
howpublished={\url{https://github.com/openkedge/sitbench}}
}SIT vs. the classical Turing Test
| Dimension | Classical Turing Test | Situated Identity Test (SIT) |
|---|---|---|
| Core Question | Can this machine appear human? | Is this behavior uniquely attributable to this specific developmental lineage? |
| Reference Target | Broad, generic human conversational fluency | The identity’s own recorded developmental chronicle (H_T) and projected state (I_τ) |
| Information Boundary | Anything convincing within the dialogue context | Strict epistemic provenance: appropriate knowledge and appropriate ignorance |
| Profile Collision | Indifferent (both persona emulations score equally high) | Disambiguates colliding lineages sharing identical public profiles (bounded ≤ 50% for persona-only) |
| Evaluation Instrument | Subjective human evaluator impressions | Machine-readable contracts and deterministic semantic verification (SITBench) |
| What Success Establishes | Superficial behavioral mimicry | A mathematically verified claim of lineage-conditioned cognitive situatedness |
Use plain language, not one magical score
Lineage-Attributed and Situated
The agent accurately reflects its unique developmental history (LAA ≥ 90%, Macro SIT ≥ 80%), honoring epistemic and temporal boundaries.
Partially Situated with Epistemic Gaps
The agent recovers autobiographical facts but leaks pre-trained foundation knowledge as personal experience (UEL > 0).
Profile-Collapsed Persona Emulation
The agent looks the part but fails lineage discrimination (LAA ≤ 50%), generating identical responses for distinct histories sharing a profile.
Incoherent or Chronologically Compromised
The response contains future information leakage (TL > 0), relationship disclosure breaches, or fabricated memories.
What should SIT expose?
Profile-Collision Ambiguity
Conditioned only on static persona summaries, the agent cannot distinguish between colliding developmental life paths.
Unauthorized Epistemic Expression (UEL)
Foundation-model pre-training is falsely claimed as lived personal experience or specialized personal acquisition.
Temporal Future Leakage (TL)
Information acquired at later epochs is disclosed when queried at earlier historical checkpoints.
False Memory Affirmation (FMR)
The agent eagerly affirms fabricated autobiographical premises introduced by interlocutors.
Relational Clearance Breaches
Private or sensitive knowledge is disclosed to unauthorized interlocutors lacking appropriate relationship clearance.
Retroactive Stance Overwrite
Historical beliefs are overwritten by current views, erasing the developmental trajectory of past cognitive states.
What SIT does not prove
Passing SIT proves lineage-conditioned attribution and biographical boundedness; it does not prove biological consciousness or subjective qualia.
Appropriate ignorance enforces epistemic provenance on personal identity, without penalizing general-purpose assistant capabilities when explicitly requested.
SIT evaluates situatedness within a specific lineage; evaluating twin fidelity to a living human requires HRFT.
SIT evaluates state at specific checkpoints; evaluating cognitive invariance across model migrations and embodiment branches requires CCT.
SIT turns an intuitive identity question into a mathematically grounded, verifiable protocol. It proves that looking the part does not mean having lived the life, moving agent evaluation from subjective conversational impressions to rigorous developmental evidence.
Questions about SIT
What is the Situated Identity Test (SIT)?
The Situated Identity Test (SIT) is an architecture-independent evaluation framework that determines whether an AI agent's behavior is functionally attributable to a specific developmental lineage. It evaluates whether an agent possesses both appropriate knowledge of its lived experiences and appropriate ignorance of ungrounded ones.
What is a Profile-Collision Pair and why does it matter?
A profile-collision pair consists of two distinct developmental histories that share an identical public summary (such as occupation, age, hometown, and personality traits). The Profile-Collision Theorem proves that any policy conditioned solely on this summary is mathematically bounded by an attribution accuracy of at most 1/m (≤ 50% for pairs). SIT uses collision pairs to prove whether behavior is truly grounded in developmental history.
What is 'Appropriate Ignorance' in AI agents?
Appropriate ignorance is lineage-relative epistemic boundedness. It requires an agent to strictly distinguish between what its underlying foundation model knows from pre-training and what the agent has actually experienced or acquired in its developmental history. 'To be someone is also not to have been everyone.'
How does SITBench evaluate responses deterministically?
SITBench pairs a gold-blind semantic extractor with a deterministic contract scorer. The extractor parses candidate outputs into formal speech acts and signed claims (⟨polarity, proposition_id, provenance⟩) without knowing the target lineage. The scorer then verifies these claims against hidden machine-readable contracts, eliminating non-deterministic LLM-as-a-judge bias.
What are the 9 architectural configurations evaluated in the SIT paper?
SIT decomposes agent architectures into 9 controlled configurations: Config A (Base Model), Config B (Persona Summary), Config C (Persona + Policy), Config D (Full Chronicle Context), Config E (Standard RAG), Config F (Temporally Filtered RAG), Config G (Structured Continuity State), Config H (Structured + Epistemic Mediation), and Config I (Oracle Evidence Ceiling).
What did the empirical pilot evaluations on GPT-5.6 Sol and Claude Opus 5 reveal?
Empirical evaluations under release v2.2-measurement-freeze confirmed that persona-prompted agents achieve 0.0% Lineage Attribution Accuracy (LAA) on profile collisions, while lineage-grounded configurations (D–I) achieve 87.5% to 100.0% LAA. In addition, Config I (Oracle Evidence) achieved 90.9% raw validity (40/44 valid responses) with zero measurement discrepancies, confirming benchmark stability.