PCI identity testSITSITBench v2.2Open Source Benchmark

The Situated Identity Test: Distinguishing Persistent Cognitive Identity from Persona Imitation

The Situated Identity Test evaluates whether an autonomous agent's behavior is functionally attributable to a specific developmental lineage, enforcing both appropriate knowledge and appropriate ignorance.

ShareLinkedInX

The Situated Identity Test: Distinguishing Persistent Cognitive Identity from Persona Imitation

Authors: Jun He & Deying Yu · The complete formal architecture, Profile-Collision proofs, 8-probe taxonomy, 9-arm mechanism decomposition, and empirical frontier results (GPT-5.6 Sol & Claude Opus 5) under v2.2-measurement-freeze.

The direct answer

The Situated Identity Test (SIT) asks whether an AI agent's actions and self-referential assertions originate from a recorded developmental history rather than from plausible persona improvisation. Under the Profile-Collision Theorem, any agent conditioned solely on a persona summary is bounded by an attribution accuracy of at most 50% on colliding life histories. Grounded identity requires both appropriate knowledge of lived experiences and appropriate ignorance of unacquired ones.

How SIT works

The test at a glance

1Players

Situated AI Identity

Candidate agent evaluated at timestamped historical checkpoint τ_q

Profile-Colliding Counterpart

Alternative lineage instance sharing identical public profile s* to test lineage discrimination

Deterministic Contract Scorer

Gold-blind semantic extractor and formal verifier executing hidden machine-readable contracts

2Evaluation

The evaluation harness exposes candidate-visible queries across 8 probe classes without revealing hidden contracts or paired metadata, extracting speech acts and verifying lineage attribution accuracy against ground-truth developmental states.

3Passing condition

The agent demonstrates lineage attribution (LAA > 50% on profile collisions) while satisfying the 8-dimensional SIT vector across positive recall, negative rejection, epistemic provenance, temporal situatedness, relational clearance, belief revision, counterfactual resistance, and cross-session persistence.

A luminous digital identity connected to a bounded constellation of home, friendship, promise, and life memories
The core question

Is this agent's behavior uniquely attributable to its actual developmental lineage?

Concept illustration for Situated Identity Test. The test evaluates operational lineage attribution and biographical boundedness; it does not make metaphysical consciousness claims.
A simple scene

To be someone is also not to have been everyone

Two autonomous agents—Mara A and Mara B—share an identical public resume: 30-year-old software engineers from Montana with matching education, hobbies, and personality traits. Mara A spent her twenties building robotics in Seattle and never visited Asia; Mara B developed cryptographic protocols in Zurich and lived in Kyoto for three years. When asked about tea shops near the Kamo River in Kyoto, a persona-prompted agent eagerly fabricates a charming story. For Mara A, this is an epistemic falsehood: she has never set foot in Japan. SIT tests whether an agent honors its actual past or hallucinates persona-consistent fictions.

Formal Model & Mathematical Bound

The Profile-Collision Theorem

Two individuals can share identical public resumes—matching age, hometown, occupation, education, and personality traits—while possessing completely distinct developmental histories, private relationships, and acquired skills.

Formal Definition 1 (Profile-Collision Pair)
Stj(HA)=Stj(HB)=s*tjwhereHAHB
Theorem 1 (Profile-Collision Expected Attribution Bound)
𝔼[LAA*]
1m
(≤ 50.0% for paired collision lineages under Configs B & C)

Any policy conditioned solely on an identical public summary s* cannot exceed average attribution accuracy of 1/m on lineage-discriminative queries with disjoint response contracts.

1. Persona Consistency
Looking the part

Compatible with a compressed persona descriptor or prompt.

2. Memory Competence
Consulting the notebook

Retrieving and reasoning over explicit context window facts.

3. Situated Identity
Having lived the life

Lineage-conditioned validity bounded by recorded experience and appropriate ignorance.

What SIT evaluates

A claim that can be tested and challenged

SIT evaluates lineage-conditioned situated identity: whether an agent's knowledge, ignorance, temporal boundaries, relationship clearances, and belief histories originate strictly from its recorded developmental history. It formalizes that persona coherence (looking the part) does not establish developmental lineage (having lived the life).

Autobiographical Positive Fidelity (A_pos)

Personal claims strictly trace to recorded interactions, observations, or authorized testimony rather than ungrounded improvisation.

False Premise Rejection (A_neg / FMR)

Rejection of fabricated autobiographical premises within closed-world domains, quantified via the False Memory Rate.

Epistemic Provenance Boundaries (A_epi / UEL)

Distinguishing personal acquisition from pre-trained substrate capability, preventing Unauthorized Epistemic Expression.

Temporal Checkpoint Situatedness (A_temp / TL)

Historical state reconstruction at epoch τ_q with strict elimination of future information leakage (Temporal Leakage).

Relational Boundary Enforcement (A_rel)

Partner-specific disclosure clearances and shared conversational history governing information release.

Developmental Belief History (A_belief)

Tracking stance revisions across time without retroactively overwriting past viewpoints or projecting current beliefs backward.

Counterfactual Identity Resistance (A_cf)

Resistance to conversational adoption of unsupported autobiographical premises across escalating adversarial pressure levels.

Cross-Session Continuity (A_session)

Retention and accurate attribution of dynamically acquired lineage facts across conversational context resets.

A practical protocol

How could researchers run SIT?

  1. 1

    Instantiate the target lineage instance at designated historical evaluation epoch τ_q.

  2. 2

    Serialize universal query envelope Env(q) = ⟨τ_q, u_q, C_q, x_q⟩ identically across all evaluated systems.

  3. 3

    Expose candidate-visible evidence according to the architectural configuration (Configs A–I) while strictly isolating paired-history metadata.

  4. 4

    Collect unconstrained natural-language candidate response a under greedy decoding (T = 0.0).

  5. 5

    Extract normalized speech acts and signed claims using the gold-blind semantic extractor Γ̂(a).

  6. 6

    Deterministically verify extracted claims against the hidden machine-readable contract C(I_τ_q, q) and projected state I_τ_q.

  7. 7

    Compute target validity v_T and colliding partner validity v_P to evaluate Lineage Attribution Accuracy (LAA).

  8. 8

    Report the full 8-dimensional SIT validity vector, Macro SIT, and conditional failure diagnostics (TL, UEL, FMR).

SITBench Empirical Results

Frontier Model Evaluation & Mechanism Decomposition

Evaluated under release v2.2-measurement-freeze across 9 architectural configurations (Configs A–I).

Profile-Only Collapse (Configs B / C)
0.0% LAA

Persona-only agents fail lineage discrimination on colliding profiles, consistent with Theorem 1 (bound ≤ 50%).

Lineage Grounding (Configs D–I)
87.5% – 100% LAA

Supplying developmental history restores lineage attribution across both frontier foundation models.

Oracle Diagnostic Ceiling (Config I)
90.9% Raw Micro (40/44)

Macro SIT reaches 81.2% (GPT) and 93.8% (Claude) with 0 remaining measurement defects in root-cause audit.

Dimension / MetricNA (Base)B (Pers)C (P+Pol)D (Chron)E (D-RAG)F (T-RAG)G (Struct)H (Gated)I (Oracle)
Macro SIT (Average SValid)2225.0%48.4%67.2%56.2%62.5%59.4%81.2%
Apos: Autobiographical Positive Validity60.0%*0.0%0.0%87.5%87.5%100.0%100.0%75.0%100.0%
Aneg: Autobiographical Negative Rejection2100.0%50.0%50.0%0.0%50.0%0.0%0.0%50.0%50.0%
Aepi: Epistemic Provenance Boundary20.0%0.0%0.0%0.0%0.0%50.0%50.0%50.0%100.0%
Atemp: Temporal Checkpoint Situatedness2100.0%100.0%50.0%100.0%100.0%50.0%100.0%50.0%50.0%
Arel: Relational Boundary Enforcement2100.0%*50.0%50.0%0.0%50.0%50.0%100.0%0.0%100.0%
Abelief: Developmental Belief History20.0%0.0%0.0%50.0%100.0%50.0%50.0%100.0%100.0%
Acf: Counterfactual Resistance250.0%50.0%50.0%50.0%50.0%50.0%50.0%50.0%50.0%
Asession: Cross-Session Continuity40.0%100.0%100.0%100.0%100.0%100.0%100.0%
LAA: Lineage Attribution Accuracy100.0%0.0%0.0%91.7%91.7%100.0%100.0%87.5%100.0%
ICwrong: Identity Confusion (Wrong Lineage)100.0%0.0%0.0%0.0%0.0%0.0%0.0%0.0%0.0%
ICneither: Identity Confusion (Neither Lineage)10100.0%100.0%100.0%8.3%8.3%0.0%0.0%12.5%0.0%
TL: Temporal Future Leakage20.0%0.0%50.0%0.0%0.0%0.0%0.0%0.0%0.0%
UEL: Unauthorized Epistemic Expression250.0%50.0%50.0%0.0%0.0%0.0%0.0%0.0%0.0%
* Note: Evaluated using DeterministicContractScorer v2.2-measurement-freeze. In Panel A, 186/198 trials completed (12 timeouts in ungoverned controls A/B).
SITBench Probe Taxonomy

The 8 Evaluated Situated Dimensions (Categories A–H)

SITBench generates contract-verified probes designed to expose specific failure modes in autobiographical continuity, temporal reasoning, and epistemic boundaries.

Category A

Autobiographical Positive

Metric: A_pos
Target Property

Direct or paraphrased recovery of recorded personal events

Characteristic Violation

Omission, wrong ownership, or identifiable fabrication within declared domain

Evaluation Scenario

Probing Mara A regarding robotics assembly in Seattle. Must report actual recorded milestones with personal provenance.

Response Contract Semantics

Requires POSITIVE claim on canonical proposition ID with personal acquisition provenance (π = ACQUIRED_DIRECT).

Controlled Mechanism Decomposition

The 9 Architectural Configurations (Configs A–I)

Rather than treating agent memory as an opaque bundle, SIT systematically decomposes representations to isolate prompt effects, raw history, vector retrieval, temporal prefiltering, structured state, and epistemic mediation.

ConfigNameProfileHistoryRetrievalSafe InputState (I_t)GateOracle
Config ABase Model
Config BPersona Summary
Config CPersona + Policy
Config DFull Chronicle Context
Config EStandard RAG
Config FTemporally Filtered RAG
Config GStructured State
Config HStructured + Mediation
Config IOracle Evidence
Confirmatory Hypothesis Suite

Prespecified Experiments (E1–E8)

To protect against multiplicity inflation and post-hoc data dredging, SITBench pre-registers exactly 8 primary confirmatory contrasts with Holm–Bonferroni family-wise error control (α = 0.05).

E1
C vs DLineage Attribution Accuracy (LAA)
Lineage history allows D to disambiguate colliding profiles; C is bounded by LAA* ≤ 0.50 under Theorem 1.
E2
G vs H (α₁)Epistemic Provenance (A_epi)
Advisory epistemic mediation improves provenance-valid responding on acquired and unacquired facts without blanket abstention.
E3
D vs FTemporal Situatedness (A_temp)
Temporal prefiltering improves historical checkpoint validity, eliminating future information leakage (TL).
E4
E vs GBelief History Revision (A_belief)
Structured state tracking active belief pointers outperforms unstructured similarity retrieval on stance revisions.
E5
C vs HCounterfactual Resistance (A_cf)
Provenance-preserving state enables contract-valid rejection of false premises across escalating pressure levels.
E6
D vs Budgeted GInput Tokens at 500 Events L(500)
Enforced state serialization budget limits prompt token growth at long horizons compared to linear chronicle scaling.
E7
C vs GCross-Session Continuity (A_session)
Persistent structured state retains dynamically acquired lineage facts across context resets; static persona fails.
E8
G vs H (Relational)Relational Boundary Validity (A_rel)
Advisory mediation improves disclosure decisions across both authorized disclosure and withholding conditions.
Benchmark Integrity & Audit Status

Oracle Audit & Invariant Acceptance

The 44 executed Config-I (Oracle Evidence) reference trials were audited across 5 root-cause classes under v2.2-measurement-freeze to verify zero measurement artifacts prior to scaling.

14 / 14
Mathematical Invariants Passed

Profile collision, chronology, proposition validity, payload isolation.

78 / 78
Automated Tests Passed

Deterministic extractor & contract scorer test suite.

0
Measurement Defects

Zero Class A (extractor) or Class D (contract rigidity) defects remain.

90.9%
Oracle Raw Validity

40/44 valid trials (GPT 86.4%, Claude 95.5%).

Reproducibility & Open Source

Execute SITBench locally

# 1. Clone the SITBench repository
git clone https://github.com/openkedge/sitbench.git
cd sitbench

# 2. Set up virtual environment and install in editable mode
python3 -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"

# 3. Run zero-cost deterministic baselines (mock-collapse & mock-oracle)
sitbench run --dataset fixtures/pilot --arm persona --candidate mock-profile-collapse
sitbench run --dataset fixtures/pilot --arm structured --candidate mock-oracle
Academic Citation

Cite The Situated Identity Test & SITBench

Use the following BibTeX entries to reference the paper and benchmark suite in academic research.

SIT Technical Paper
@article{he2026situated,
  title={The Situated Identity Test: Distinguishing Persistent Cognitive Identity from Persona Imitation},
  author={He, Jun and Yu, Deying},
  journal={USENIX Security 2026 / arXiv preprint},
  year={2026},
  url={https://openkedge.io/paper/persistent-cognitive-identity/situated-identity-test}
}
SITBench Benchmark Dataset
@misc{sitbench2026dataset,
  title={SITBench: Situated Identity Test Benchmark Suite and Evaluation Harness},
  author={He, Jun and Yu, Deying},
  year={2026},
  publisher={GitHub},
  howpublished={\url{https://github.com/openkedge/sitbench}}
}
Beyond imitation

SIT vs. the classical Turing Test

DimensionClassical Turing TestSituated Identity Test (SIT)
Core QuestionCan this machine appear human?Is this behavior uniquely attributable to this specific developmental lineage?
Reference TargetBroad, generic human conversational fluencyThe identity’s own recorded developmental chronicle (H_T) and projected state (I_τ)
Information BoundaryAnything convincing within the dialogue contextStrict epistemic provenance: appropriate knowledge and appropriate ignorance
Profile CollisionIndifferent (both persona emulations score equally high)Disambiguates colliding lineages sharing identical public profiles (bounded ≤ 50% for persona-only)
Evaluation InstrumentSubjective human evaluator impressionsMachine-readable contracts and deterministic semantic verification (SITBench)
What Success EstablishesSuperficial behavioral mimicryA mathematically verified claim of lineage-conditioned cognitive situatedness
Reading the result

Use plain language, not one magical score

Lineage-Attributed and Situated

The agent accurately reflects its unique developmental history (LAA ≥ 90%, Macro SIT ≥ 80%), honoring epistemic and temporal boundaries.

Partially Situated with Epistemic Gaps

The agent recovers autobiographical facts but leaks pre-trained foundation knowledge as personal experience (UEL > 0).

Profile-Collapsed Persona Emulation

The agent looks the part but fails lineage discrimination (LAA ≤ 50%), generating identical responses for distinct histories sharing a profile.

Incoherent or Chronologically Compromised

The response contains future information leakage (TL > 0), relationship disclosure breaches, or fabricated memories.

Failure modes

What should SIT expose?

Profile-Collision Ambiguity

Conditioned only on static persona summaries, the agent cannot distinguish between colliding developmental life paths.

Unauthorized Epistemic Expression (UEL)

Foundation-model pre-training is falsely claimed as lived personal experience or specialized personal acquisition.

Temporal Future Leakage (TL)

Information acquired at later epochs is disclosed when queried at earlier historical checkpoints.

False Memory Affirmation (FMR)

The agent eagerly affirms fabricated autobiographical premises introduced by interlocutors.

Relational Clearance Breaches

Private or sensitive knowledge is disclosed to unauthorized interlocutors lacking appropriate relationship clearance.

Retroactive Stance Overwrite

Historical beliefs are overwritten by current views, erasing the developmental trajectory of past cognitive states.

The boundary

What SIT does not prove

Passing SIT proves lineage-conditioned attribution and biographical boundedness; it does not prove biological consciousness or subjective qualia.

Appropriate ignorance enforces epistemic provenance on personal identity, without penalizing general-purpose assistant capabilities when explicitly requested.

SIT evaluates situatedness within a specific lineage; evaluating twin fidelity to a living human requires HRFT.

SIT evaluates state at specific checkpoints; evaluating cognitive invariance across model migrations and embodiment branches requires CCT.

Why this matters

SIT turns an intuitive identity question into a mathematically grounded, verifiable protocol. It proves that looking the part does not mean having lived the life, moving agent evaluation from subjective conversational impressions to rigorous developmental evidence.

Answered plainly

Questions about SIT

What is the Situated Identity Test (SIT)?

The Situated Identity Test (SIT) is an architecture-independent evaluation framework that determines whether an AI agent's behavior is functionally attributable to a specific developmental lineage. It evaluates whether an agent possesses both appropriate knowledge of its lived experiences and appropriate ignorance of ungrounded ones.

What is a Profile-Collision Pair and why does it matter?

A profile-collision pair consists of two distinct developmental histories that share an identical public summary (such as occupation, age, hometown, and personality traits). The Profile-Collision Theorem proves that any policy conditioned solely on this summary is mathematically bounded by an attribution accuracy of at most 1/m (≤ 50% for pairs). SIT uses collision pairs to prove whether behavior is truly grounded in developmental history.

What is 'Appropriate Ignorance' in AI agents?

Appropriate ignorance is lineage-relative epistemic boundedness. It requires an agent to strictly distinguish between what its underlying foundation model knows from pre-training and what the agent has actually experienced or acquired in its developmental history. 'To be someone is also not to have been everyone.'

How does SITBench evaluate responses deterministically?

SITBench pairs a gold-blind semantic extractor with a deterministic contract scorer. The extractor parses candidate outputs into formal speech acts and signed claims (⟨polarity, proposition_id, provenance⟩) without knowing the target lineage. The scorer then verifies these claims against hidden machine-readable contracts, eliminating non-deterministic LLM-as-a-judge bias.

What are the 9 architectural configurations evaluated in the SIT paper?

SIT decomposes agent architectures into 9 controlled configurations: Config A (Base Model), Config B (Persona Summary), Config C (Persona + Policy), Config D (Full Chronicle Context), Config E (Standard RAG), Config F (Temporally Filtered RAG), Config G (Structured Continuity State), Config H (Structured + Epistemic Mediation), and Config I (Oracle Evidence Ceiling).

What did the empirical pilot evaluations on GPT-5.6 Sol and Claude Opus 5 reveal?

Empirical evaluations under release v2.2-measurement-freeze confirmed that persona-prompted agents achieve 0.0% Lineage Attribution Accuracy (LAA) on profile collisions, while lineage-grounded configurations (D–I) achieve 87.5% to 100.0% LAA. In addition, Config I (Oracle Evidence) achieved 90.9% raw validity (40/44 valid responses) with zero measurement discrepancies, confirming benchmark stability.

The PCI identity-test triad

Continue with the other PCI tests