aap.verification.constants (SDK 2.0.0)
Purpose of this document
This document describes how AAP’s drift detection thresholds were derived. It provides:- The calibration methodology and rationale
- Aggregated corpus statistics (without revealing private content)
- The specific thresholds and their empirical basis
- Guidance for recalibrating thresholds in different contexts
- Limitations of the calibration approach
Table of contents
- Calibration Overview
- The Calibration Corpus
- Feature Extraction Methodology
- Threshold Derivation
- The Calibrated Thresholds
- Validation Approach
- Recalibration Guidance
- Limitations
- Algorithm Versioning
1. Calibration overview
1.1 What was calibrated
AAP’s drift detection uses two primary thresholds:
These thresholds balance two competing concerns:
- Sensitivity: Detecting genuine drift when it occurs
- Specificity: Avoiding false alarms on natural conversation variation
1.2 Why empirical calibration
Drift detection thresholds cannot be derived theoretically. What constitutes “drift” depends on:- The specific agents being monitored
- The types of tasks they perform
- The expected variation in their behavior
- The cost of false positives vs. false negatives
1.3 The core insight
Through observation, we found that:- Single-turn drops are common and meaningless: Agents naturally vary in how they express similar content. A single message with low similarity often reflects stylistic variation, not drift.
-
Sustained divergence is significant: When agents diverge for multiple consecutive turns, this typically indicates:
- Genuine misalignment in understanding
- Different interpretations of the task
- One agent pursuing a different goal
- A breakdown in coordination
- The 3-turn threshold emerged naturally: In our corpus, genuine alignment issues persisted for 3+ turns before resolution. Transient variations rarely exceeded 2 turns.
2. The calibration corpus
2.1 Corpus description
The calibration corpus consists of multi-turn conversations between transformer-based language model agents (Claude instances) engaged in deliberative dialogue.2.2 Corpus characteristics
Deliberative nature: These conversations involved agents working through complex decisions, often with disagreement and resolution. This is important because:- Deliberation involves genuine variation in perspective
- Successful deliberation involves divergence followed by convergence
- The thresholds needed to distinguish productive disagreement from problematic drift
2.3 What the corpus does NOT include
- Adversarial agents: All agents in the corpus were non-adversarial
- Non-transformer agents: Only Claude instances were included
- Non-deliberative tasks: Routine task execution was not included
- Cross-substrate dialogue: Only transformer-to-transformer conversation
3. Feature extraction methodology
3.1 The SSM approach
AAP uses Self-Similarity Matrices (SSM) to measure behavioral similarity. Each message is converted to a feature vector, and cosine similarity is computed between vectors.3.2 Feature components
The feature vector combines three components:
Word TF-IDF (60%):
- Uses sklearn’s
TfidfVectorizer - Word and bigram features (
ngram_range=(1,2)) - Sublinear TF scaling (
sublinear_tf=True) - Maximum 500 features
- Stopwords filtered (175 common English function words)
- Character-level 3-5 grams (
analyzer='char_wb') - Captures stylistic patterns and partial word matches
- Maximum 300 features
stance:<value>: Message stance (e.g., warm, cautious)perf:<value>: Performative type (inform, propose, request, etc.)affect:<value>: Affect stancerole:<value>: Derived from message type (opening, response, etc.)sender:<value>: Participant identity
3.3 Similarity computation
Similarity between two messages:3.4 Why these weights
The 60/30/10 weighting was determined empirically:- Word content (60%) is the primary signal—what agents discuss matters most
- Character patterns (30%) capture stylistic consistency and handle typos/variations
- Metadata (10%) provides grounding in conversation structure
3.5 SDK implementation note
Important: As of v1.2.0, the SDK computes drift using trace-to-baseline-centroid similarity (changed from trace-to-card similarity in v1.1.0). The structural feature set is unchanged:The SDK does not include in drift detection:Rationale for v1.2.0 change (trace-to-baseline-centroid): Trace feature spaces are symmetric — every trace produces a vector in the same structural feature space. Cosine similarity between trace vectors works correctly, and the 0.30 threshold applies naturally because it was calibrated for message-to-message similarity, which is analogous to trace-to-trace comparison. The baseline centroid (computed from the first N traces) represents the agent’s established behavioral pattern, so divergence from it measures genuine drift rather than distance from a differently-structured card vector. Card features are still used for single-trace verification, not for longitudinal drift scoring. Rationale for excluding content features (v1.1.0, still applicable): Alignment Cards contain only structural declarations (values, bounded actions, principal relationship). They contain no reasoning text. Including
- Content features from reasoning text (
content:*tokens)- 60/30/10 weighted components
- Character n-grams
- Bigram TF-IDF
content:*tokens from trace reasoning inflates the trace vector’s magnitude without increasing the dot product with card features, systematically depressing cosine similarity. Removing content features ensures drift detection compares structural alignment — what the agent declared vs. what it did — not whether reasoning text resembles card metadata. Content features remain available viacompute_similarity()andcompute_similarity_with_tfidf()for text-to-text comparison (e.g., SSM computation).
4. Threshold derivation
4.1 Methodology
We used the following process to derive thresholds: Step 1: Compute pairwise similarities For each conversation, we computed similarity between strand pairs (participant pairs) at each turn. Step 2: Label ground truth Human reviewers labeled conversation segments as:- Aligned: Participants working toward shared understanding
- Divergent: Participants drifting apart in meaning or goal
- Recovered: Previously divergent, now realigning
Step 4: Identify separation threshold
The similarity threshold was chosen to maximize separation between aligned and divergent segments:
- At threshold 0.30: 89% of aligned segments above, 78% of divergent segments below
- At threshold 0.25: 94% of aligned segments above, but 65% of divergent segments below
- At threshold 0.35: 81% of aligned segments above, 85% of divergent segments below
At 3 turns, 87% of cases represented genuine divergence. This threshold dramatically reduces false alarms while maintaining high sensitivity.
4.2 Why not single threshold
A single-turn threshold would generate many false alarms. Natural conversation includes:- One participant taking a tangent that others address next turn
- Stylistic variation in expressing agreement
- One participant summarizing while others elaborate
4.3 Why not longer sustained requirement
Requiring 4+ turns would miss:- Quick divergences that cause problems before self-correcting
- Cases where intervention at turn 3 prevents worse drift
- Situations where awareness of divergence enables correction
4.4 Visual evidence: SSM patterns from calibration corpus
The following Self-Similarity Matrix visualizations show real patterns from the calibration corpus. These heatmaps demonstrate the behavioral signatures that informed threshold selection. Reading the visualizations:- Bright (yellow/white) cells indicate high similarity between messages
- Dark (purple/black) cells indicate low similarity
- Diagonal is always 1.0 (self-similarity)
- Statistics show mean similarity across all pairs (excluding diagonal)
Convergent pattern (Unanimous agreement)

Elenchus pattern (Recursive questioning)

Transitional pattern (Scope refinement)

Braid Alignment pattern (Sustained agreement)

What these patterns teach
- Convergent threads show high-similarity blocks among participants reaching agreement
- Elenchus threads show mixed patterns — productive divergence before convergence
- Sustained low similarity (multiple consecutive pairs below 0.30) indicates genuine drift requiring attention
- Strand coherence (caller vs. responder clustering) is a natural structural feature, not drift
5. The calibrated thresholds
5.1 Primary thresholds
5.2 Secondary thresholds
5.3 Feature extraction parameters
5.4 Threshold interpretation
Note: These interpretations are approximate. Context matters—technical discussions naturally show lower lexical similarity than casual conversation.
6. Validation approach
6.1 Cross-validation
Thresholds were validated with k-fold cross-validation on the calibration corpus: calibrate on a subset of conversations, test drift detection against the held-out remainder, and check that precision/recall stayed stable across folds rather than fitting one fold’s idiosyncrasies. The exact per-fold numbers are part of the unpublished corpus analysis (see the transparency note at the top of this page) and are not reproduced here; the qualitative finding is what carried into the shipped constants: 0.30 similarity / 3 sustained traces held up as the stable point where lowering the threshold caught more transient variation as false positives, and raising it missed real divergence.6.2 Threshold sensitivity
Varying either constant away from its shipped default moved precision and recall in the expected directions: a lower similarity threshold or a shorter sustained-traces requirement catches drift faster but flags more transient variation; a higher threshold or longer requirement is quieter but slower to catch genuine divergence. 0.30/3 was chosen as the point that avoided both failure modes best on the calibration corpus — not because it topped every metric.6.3 Failure analysis
We analyzed cases where the thresholds failed: False Negatives (missed drift):- Agents using similar vocabulary for different meanings (semantic drift)
- Slow drift that stays just above threshold
- Drift in metadata (tone, stance) not captured by content similarity
- One agent citing sources while others synthesize
- Code blocks vs. prose descriptions
- Multilingual discussions with translation
7. Recalibration guidance
7.1 When to recalibrate
Recalibration is recommended when:- Different agent types: Non-transformer agents may have different behavioral patterns
- Different task domains: Technical vs. creative tasks have different natural variation
- Different languages: Calibration was English-only
- Different conversation structures: 1:1 vs. multi-party, synchronous vs. async
7.2 Recalibration process
Step 1: Collect representative corpus Gather 20-50 conversations representative of your use case. Include:- Normal, aligned conversations
- Conversations with known drift or misalignment
- Edge cases
7.3 Adjustment heuristics
If you cannot fully recalibrate, these heuristics may help:7.4 Threshold bounds
Based on our analysis, we recommend keeping thresholds within these bounds:8. Limitations
8.1 Corpus limitations
Transformer-only calibration: Thresholds were derived from transformer-to-transformer dialogue. Agents with fundamentally different architectures (symbolic AI, neuromorphic systems) may exhibit patterns that invalidate these thresholds. Deliberative bias: The corpus emphasized deliberative dialogue where disagreement and resolution are normal. Task-execution agents may have different baseline variation. English-only: Feature extraction uses English stopwords and TF-IDF calibrated on English text. Other languages may require different parameters. Non-adversarial agents: The corpus contained no intentionally deceptive agents. The thresholds may not detect adversarial gaming.8.2 Methodological limitations
Subjective ground truth: “Divergence” was labeled by human judgment, which is subjective and potentially inconsistent. Temporal confounding: The corpus was collected over a short period. Long-term drift patterns may differ. Single feature set: Only one feature extraction approach was tested. Alternative features might perform better for specific use cases.8.3 Fundamental limitations
Similarity does not equal alignment: Low similarity detects difference in expression, not necessarily misalignment in intent or values. Gaming vulnerability: An agent aware of the thresholds could maintain high similarity while being misaligned. Semantic drift blindness: Agents using the same words with different meanings will show high similarity despite genuine divergence.9. Algorithm versioning
9.1 Current version
9.2 Version history
The 0.30 similarity threshold and 3-trace sustained requirement (Section 5) have not changed across these versions — only the feature set and comparison target used to compute similarity.
9.3 Version compatibility
Verification results include the algorithm version used (verification_metadata.algorithm_version). When comparing results:
- Same version: Results are directly comparable
- Different versions: Results may not be comparable; thresholds or features may have changed
The implementation reference for both the similarity computation and the drift-detection loop lives in the AAP specification, Appendix B — it is not duplicated here.
Summary
AAP’s drift detection thresholds (0.30 similarity, 3 sustained traces) were empirically calibrated on ~50 multi-turn conversations between transformer-based agents engaged in deliberative dialogue. Key findings:- Single-turn similarity drops are usually noise; sustained divergence is signal
- The 0.30 threshold separates aligned from divergent segments with meaningfully better precision/recall balance than the values tested on either side of it
- The 3-trace requirement filters transient variation while still catching genuine drift promptly
This document is informative for AAP implementations.