chore(pi): add session affinity and cache diagnostics - #336
Open
dgokeeffe wants to merge 9 commits into
Open
Conversation
…directing HOME Pi honors the PI_CODING_AGENT_DIR env var to resolve its config directory (~/.pi/agent), so redirecting /Users/david.okeeffe to APP_DIR/pi-home was unnecessary. The HOME redirect broke macOS keychain default resolution under ucode: the Security framework looks for the login keychain under the redirected HOME, finds none, and security default-keychain returns 'A default keychain could not be found'. As a result gh auth, the git credential helper, and any keychain-backed tool failed inside pi. Setting PI_CODING_AGENT_DIR to the existing PI_CONFIG_DIR preserves config isolation (models.json/settings.json/sessions still land under APP_DIR/pi-home/.pi/agent) while leaving /Users/david.okeeffe as the user's real home, so the login keychain stays discoverable. Tests: the two pi e2e sites monkeypatched PI_UCODE_HOME/PI_CONFIG_PATH to redirect pi at a tmp home. They now also patch PI_CONFIG_DIR (read by build_runtime_env) and PI_SETTINGS_PATH/PI_SETTINGS_BACKUP_PATH (previously masked because the HOME redirect made pi read settings from the un-patched real APP_DIR path).
Expose the GLM and Kimi coding-model cohort through Pi and OpenCode with shared token limits and reasoning metadata. Keep unsupported chat models out of discovery, including Inkling until gateway issue databricks#215 is fixed, and retain the GPT-OSS Responses API routing guard.
Centralize Claude family/version parsing so Pi metadata, adaptive-thinking compatibility, and Claude Code's [1m] selector cannot drift. Cover Sonnet 4.5, Opus 4.6, future major versions, Fable fallback, and prefixed model IDs.
Configure the Databricks OpenAI Responses provider alongside the validated GLM/Kimi provider, and fall back to foundation-model serving endpoints when UC model services are unavailable.
…/consolidated-pi-opencode-upstream # Conflicts: # tests/test_e2e.py
`_pi_gpt_model_entry` declared `reasoning: True` without an off-state, so for
the thinking-off case Pi's Responses builder fell back to
`reasoning: {effort: "none"}` (pi-ai openai-responses.js, the
`thinkingLevelMap?.off !== null` branch). `"none"` is only valid on gpt-5.1+,
so every request to gpt-5, gpt-5-mini, gpt-5-nano and gpt-5-5-pro was rejected:
BAD_REQUEST: Unsupported value: 'none' is not supported with the 'gpt-5'
model. Supported values are: 'minimal', 'low', 'medium', and 'high'.
Setting `thinkingLevelMap: {"off": None}` makes Pi omit `reasoning` entirely,
which the gateway accepts for all 14 codex ids. Verified against
/ai-gateway/codex/v1/responses: effort="none" 400s on gpt-5/-mini/-nano/-5-5-pro
and 200s on gpt-5-1..-5-6; omitting `reasoning` is 200 everywhere.
`{"off": "minimal"}` was rejected as an alternative because gpt-5-5-pro 400s on
it too. Same pattern already used for the Gemini 3.x entries.
The rest of Pi's Responses payload was bisected against the gateway and is
fine: store:false, prompt_cache_key, prompt_cache_retention:"24h",
prompt_cache_options, include:["reasoning.encrypted_content"], developer role,
flat tool schemas, and the session_id / x-client-request-id affinity headers.
Regression was hard to spot because the gateway returns
{"error_code","message"} rather than OpenAI's {"error":...}, so Pi's
error-body.js recovery no-ops and every 400 renders as
"OpenAI API error (400): 400 status code (no body)". Reported upstream as
earendil-works/pi#7748.
Refs databricks#286
…i-opencode-upstream # Conflicts: # src/ucode/agents/pi.py # src/ucode/databricks.py
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Issue and stack
Closes #331.
Depends on #217, #333, #223, #239, and #243, but is independent of #334/#335. Review only the session-affinity/diagnostic range relative to the pushed integrated base:
dgokeeffe/ucode@review/integrated-consolidation-base...review/pi-cache-affinity-independent
Equivalent immutable range:
168038b...e672d4e. Do not merge before the prerequisite stack. The only production behavior in this unit is Pi session affinity; it does not force or claim short GPT-subagent retention.Unit 07 — Pi cache/session-affinity diagnostics
Objective and user-visible behavior
Enable Pi's production-safe session-affinity headers so a conversation remains on one UAG destination when traffic splitting is configured. Add sanitized, explicit diagnostic tools/tests for prompt-cache investigations without asserting a backend retention guarantee.
Candidate:
review/pi-cache-affinity-independent(e672d4e), independent incremental basedev(168038b).Exact scope
Production: one Pi Claude-provider compat flag in
src/ucode/agents/pi.py.Diagnostics:
scripts/prompt_cache_repro.py,scripts/prompt_cache_http_trace.py.Tests: one Pi compat assertion plus
tests/test_prompt_cache_repro.py,tests/test_prompt_cache_http_trace.py.Non-goals: forcing
PI_CACHE_RETENTION, installing a child/subagent cache extension, claiming short GPT-subagent retention, MLflow discovery/proxying.Before / after reproduction
Focused validation:
Full suite:
uv run --frozen pytest -q # 2 failed, 1751 passed, 36 skippedThe installed Claude capture and preferred-Pi-model state assertion reproduce on integrated
dev; neither is touched by this six-file range.Retained cache investigation evidence
prompt_cache_keyandprompt_cache_retention: "24h".24h; invalid values reportin_memoryand24has supported values.in_memory; therefore this unit makes no short-retention claim for GPT subagents.The trace script prints sanitized request/response JSON, never prints authorization headers, redacts an echoed exact token, and persists only a cache key, assistant text, and usage metadata. Tests verify omission vs
24hvsin_memory, stable prefixes, cache usage extraction, phase/state behavior, malformed saved-state rejection without traceback, and redaction.Impact map
Rollback and residual risk
Revert the two unit-07 commits (
945b0c2,e672d4e); session affinity and diagnostic scripts disappear with no state migration. Residual risk remains backend cache eviction/forwarding/prefix behavior; the scripts diagnose but do not promise retention.Hygiene
git diff --check dev..review/pi-cache-affinity-independentpasses. Six-file range; no MLflow production files, generated outputs/state files, credentials,uv.lock,.pi-subagents/, orgoal.md; no merge markers or unresolved index entries.