fix(opencode): route GPT models and discover OSS fallbacks - #239
fix(opencode): route GPT models and discover OSS fallbacks#239dgokeeffe wants to merge 4 commits into
Conversation
bc65755 to
b53813f
Compare
|
This stack has been rebuilt on current main with one clean commit per layer. The incremental OpenCode diff is linked in the description; the rejected proxy and broad OSS cohort are gone. Local focused validation is 480 passing with Ruff clean. A maintainer CI approval/review is now the remaining gate. |
b53813f to
3f1c227
Compare
|
Post-review update: OpenCode GPT entries now carry shared context/output limits and prefer the newest eligible GPT; foundation endpoint parsing is defensive against malformed records; CLI fallback tests are deterministic; and the live e2e fixture now populates OpenAI/OSS models and isolates Pi settings paths. Round-3 reviewers found no production blocker; final validation is 588 passed/29 skipped focused and 1062 passed/36 skipped excluding environment-sensitive installed-agent capture tests. |
82e6f84 to
48a91ba
Compare
|
Rebased this PR stack onto current Hard stack dependencyThis is 3 of 3 and must land after #223:
This is a code dependency, not just a review preference: #239's The helper and its tests now live in this PR with their only consumer ( If #217/#223 are squash-merged, please rebase this branch onto the updated Validation
This PR adds lower-level OpenCode GPT routing, but managed-config OpenCode GPT mapping and admin allowlist intersection are deliberately deferred to #290. |
Expose the GLM and Kimi coding-model cohort through Pi and OpenCode with shared token limits and reasoning metadata. Keep unsupported chat models out of discovery, including Inkling until gateway issue databricks#215 is fixed, and retain the GPT-OSS Responses API routing guard.
Centralize Claude family/version parsing so Pi metadata, adaptive-thinking compatibility, and Claude Code's [1m] selector cannot drift. Cover Sonnet 4.5, Opus 4.6, future major versions, Fable fallback, and prefixed model IDs.
`_pi_gpt_model_entry` declared `reasoning: True` without an off-state, so for
the thinking-off case Pi's Responses builder fell back to
`reasoning: {effort: "none"}` (pi-ai openai-responses.js, the
`thinkingLevelMap?.off !== null` branch). `"none"` is only valid on gpt-5.1+,
so every request to gpt-5, gpt-5-mini, gpt-5-nano and gpt-5-5-pro was rejected:
BAD_REQUEST: Unsupported value: 'none' is not supported with the 'gpt-5'
model. Supported values are: 'minimal', 'low', 'medium', and 'high'.
Setting `thinkingLevelMap: {"off": None}` makes Pi omit `reasoning` entirely,
which the gateway accepts for all 14 codex ids. Verified against
/ai-gateway/codex/v1/responses: effort="none" 400s on gpt-5/-mini/-nano/-5-5-pro
and 200s on gpt-5-1..-5-6; omitting `reasoning` is 200 everywhere.
`{"off": "minimal"}` was rejected as an alternative because gpt-5-5-pro 400s on
it too. Same pattern already used for the Gemini 3.x entries.
The rest of Pi's Responses payload was bisected against the gateway and is
fine: store:false, prompt_cache_key, prompt_cache_retention:"24h",
prompt_cache_options, include:["reasoning.encrypted_content"], developer role,
flat tool schemas, and the session_id / x-client-request-id affinity headers.
Regression was hard to spot because the gateway returns
{"error_code","message"} rather than OpenAI's {"error":...}, so Pi's
error-body.js recovery no-ops and every 400 renders as
"OpenAI API error (400): 400 status code (no body)". Reported upstream as
earendil-works/pi#7748.
Refs databricks#286
Configure the Databricks OpenAI Responses provider alongside the validated GLM/Kimi provider, and fall back to foundation-model serving endpoints when UC model services are unavailable.
48a91ba to
816588f
Compare
Issue and stack
Closes #85.
Depends on #217, #333, and #223. Review only the OpenCode/discovery incremental range:
dgokeeffe/ucode@94fe107...816588f
Do not merge before #223. Once prerequisites merge, the visible PR diff collapses to the eight-file OpenCode unit.
Unit 04 — OpenCode GPT routing and OSS fallback discovery
Objective and user-visible behavior
Route GPT/Codex models through OpenCode's Databricks Responses provider and discover validated OSS serving endpoints when UC model services do not expose them.
Candidate:
review/opencode-routing(816588f), mandatory PR targetreview/pi-gpt-reasoning(94fe107), existing PR #239.Exact scope
Production:
src/ucode/agents/opencode.py,src/ucode/cli.py,src/ucode/databricks.py.Tests/fixtures:
tests/conftest.py,tests/test_agent_opencode.py,tests/test_cli.py,tests/test_databricks.py,tests/test_e2e.py.Non-goals: Pi behavior, context policy, dynamic OSS capability metadata, MLflow repair proxy, cache behavior.
Before / after reproduction
Focused validation:
Full suite:
uv run --frozen pytest -q # 2 failed, 1649 passed, 36 skippedFailures are installed Claude/Pi User-Agent capture tests. The branch predates the independent config-dir unit; integrated
devpasses the Pi capture test and independently reproduces the Claude capture failure.Live/discovery evidence
No token-bearing output is retained. Fixture coverage reproduces UC-first discovery, serving-endpoint fallback, dialect filtering, validated GLM/Kimi allowlisting, GPT provider qualification, default routing, and per-model E2E configuration with workspace-gated skips.
Impact map
databricks-openaiResponses and existingdatabricks-ossMLflow provider.Rollback and residual risk
Revert
816588f; OpenCode loses GPT routing and serving-endpoint OSS fallback but existing native-family behavior remains. Residual risk is future endpoint metadata drift; unit 05 adds structured capability validation separately.Hygiene
git diff --check review/pi-gpt-reasoning..review/opencode-routingpasses. PR #239 must target unit 02; its target-relative diff is the review contract. No generated files, credentials,uv.lock,.pi-subagents/, orgoal.md; no merge markers or unresolved index entries.