Skip to content

Latest commit

 

History

16 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

cross-episode-memory-client

Client-side pieces for cross-episode memory experiments against the feature/versionable-tables branch of Self-Reflective APIs. The server-side contract (cache_hint, table, table_versions, admin reload) is described in that branch's HOWTO_CLIENT_MEMORY.md; this repository consumes the contract and adds nothing server-side. The API repository is imported, never modified; table files touched during a run are backed up and restored.

Pieces

  • gen_phase1_stream.py -- authors the 24-episode recurrence-structured stream (10 celiac, 10 incompatible, 4 scaling controls) as an ordered list of (task, table_snapshot) pairs. Drift pattern: single table swap, celiac_brands 1.0.0 -> 2.0.0 after episode index 15. Descriptions and success criteria carry no fix values.
  • run_phase1_arms.py -- the four memory arms and the harness. A0 none, A1 naive (stores everything, ignores version stamps), A2 governed (stores only server-marked cacheable fixes, drops an entry when a response's version stamp no longer matches the stamp it was learned under), A3 budget-elastic (governed admission plus a byte budget that tightens mid-stream and recovers; the CrystalMem case, arXiv:2608.00303). The harness replays the stream and performs drift by overwriting the table file and calling the admin reload endpoint. Memory clients read cache_hint, table, and table_versions from live responses only -- they never see the stream's snapshot.
  • analyze_phase1.py -- per-arm summary: tokens per success, first-try rate, memory events, post-drift window.
  • tasks_phase1_stream.json -- the generated stream (regenerable).

Phase 2, against the feature/drift-snapshots-and-table-diff branch:

  • gen_phase2_streams.py -- re-emits the same 24 tasks under two drift patterns: v2 (mild -- one flour brand removed, oats unchanged) and v4 (severe -- every flour and oats brand rotated).
  • run_phase2_arms.py -- the Phase 1 arms plus A2D, a diff-aware governed store. The reload endpoint on this branch returns a structured diff; the harness hands it to the store, which acts on it only when a response's version stamp exposes the change, so A2 vs A2D isolates invalidation granularity under identical detection. Also keys relearned entries under the raw task's ingredient term (Phase 1 keyed them under the failing request's stale brand string, which cost one extra cold retry per drift).
  • analyze_phase2.py, tasks_phase2_v2.json, tasks_phase2_v4.json.

Phase 3, same server branch as Phase 2:

  • Second-model replication: run_phase2_arms.py reruns the Phase 2 patterns under a different PILOT_MODEL. Every live run now starts with a provider preflight -- one production-shaped probe whose requested vs. served model and reported token usage go into the results file's config block. Gateways can alias one model name to another and the accounting behind usage numbers can change between runs, so each results file carries the evidence needed to interpret its own token columns. The results filename carries the model.
  • gen_phase3_streams.py -- the multi-drift ladder: the Phase 1/2 task list repeated twice under four drift events that walk celiac_brands through every snapshot on the branch. The events alternate which half of the table survives (oats survive the flour rotation, flour survives the oats rotation), so a granularity dividend must show on both sides before it can be credited to eviction granularity itself. Streams may carry drift_events and a results_prefix; result rows gain a drift_epoch column counting the drift events seen so far.
  • analyze_phase3.py -- window summaries: tokens per success by drift epoch, the celiac-only view of each window, and the first post-drift celiac episode of every epoch (the stale-attempt slot).
  • --hint-style -- the second model exposed an instruction conflict the first model had been quietly absorbing: the base system prompt says to use the Input Data exactly as provided on the first attempt, and the Phase 2 memory block (append) asks for the opposite whenever memory matches. claude-haiku-4-5 sided with the memory block; claude-sonnet-4-6 sided with the base instruction, which silently zeroed the realized value of valid memory. The amendment wording (now the default) presents memory as validated amendments to the Input Data, so the two instructions stop competing. Runs record their wording in the config block.

Acme billing, against main at v0.2.0 (drift axis 5: a second domain):

  • gen_acme_streams.py -- authors tasks_acme_sanity.json (24 episodes, one drift) and tasks_acme_m1.json (48 episodes, three drifts) over one 12-task epoch block repeated verbatim. The drift axis is active_csm_codes: plan-to-promo-code rows whose values are unguessable, rotate periodically, and surface only in recovery_feedback -- the same memory class as the celiac brands. The snapshots alternate which row survives (the partner code rotates at v2, the enterprise-pilot code at v3, both at v4), mirroring the recipe ladder so granularity numbers are comparable across domains.
  • run_acme_arms.py -- the same arms with an Acme-vocabulary extractor. Entry keys come from raw task fields (plan id, tier, token, promo code), never from a drifting value. Dict-shaped tables surface drift as changed keys, and removed dict rows are identified by key: billing values are low-entropy vocabulary ("annual", "monthly"), and value matching would evict entries from unrelated rows that share a word. The harness also derives a companion promo_eligibility update at each rotation (base table plus the codes the csm table currently serves, version bumped), because the rotated-in code would otherwise be rejected as unknown by the promo-stacking policy; every touched file is restored after the run. Refund episodes are deliberately absent: a refund needs a live PaymentIntent id from a prior charge, so the refund stream will provision that charge at execution time.
  • acme_compat/ -- a minimal stand-in for the private stripe_api.sandbox helper the upstream app imports at startup (configure() sets the SDK key from STRIPE_API_KEY and refuses live-mode keys; normalize_error() flattens a StripeError). The runner prepends it to the server's PYTHONPATH, so the upstream tree runs unmodified from a clean clone.
  • analyze_acme.py -- the Phase 3 window summaries plus a stable-policy view (funding/enterprise/promo first-try by epoch).

Setup

Clone the API repository next to the scripts and check out the branch for the phase you are running (feature/versionable-tables for Phase 1, feature/drift-snapshots-and-table-diff for Phase 2):

git clone --branch feature/versionable-tables https://github.com/arquicanedo/self-reflective-apis

Use the same Python environment that runs the API repository's own experiments (Python 3.12; fastapi, uvicorn, anthropic, langchain). Model access is read from the standard provider environment variables; the scripts hardcode no endpoint and no key.

Acme runs use main at tag v0.2.0 and additionally need the stripe package and a STRIPE_API_KEY environment variable holding a Stripe test-mode secret key (sk_test_...): the Acme app forwards policy-clean requests to Stripe's real test API using documented pm_card_* fixtures, so no real money is involved and no PaymentMethods are created.

Run

python run_phase1_arms.py --mock     # scripted agent, no LLM calls: validates the full plumbing
python run_phase1_arms.py            # live arms (claude-haiku-4-5 by default; override with PILOT_MODEL)
python analyze_phase1.py             # summarize the latest results file

python run_phase2_arms.py --mock --pattern v2    # Phase 2 plumbing, mild drift
python run_phase2_arms.py --mock --pattern v4    # Phase 2 plumbing, severe drift
python run_phase2_arms.py --pattern v2           # live five arms
python run_phase2_arms.py --pattern v4
python analyze_phase2.py

python run_phase2_arms.py --mock --stream tasks_phase3_m1.json   # multi-drift plumbing
python run_phase2_arms.py --stream tasks_phase3_m1.json --arms A0,A1,A2,A2D
python analyze_phase3.py

python gen_acme_streams.py
python run_acme_arms.py --mock --stream tasks_acme_sanity.json   # full plumbing incl. the Stripe test call
python run_acme_arms.py --stream tasks_acme_m1.json              # live four arms
python analyze_acme.py

One contract property worth knowing

Invalidation is detect-on-response: a client only sees a new version stamp after its first attempt is already out, so a governed store pays exactly one stale attempt after a table bump before it drops the entry and relearns.

About

Client-side memory arms for the versionable-tables branch of Self-Reflective APIs

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages