A proxy that turns your GitHub Copilot subscription into an OpenAI and Anthropic compatible API. Use it to power Claude Code, Cursor, or any tool that speaks the OpenAI Chat Completions, OpenAI Responses, or Anthropic Messages protocol.
Warning
Reverse-engineered, unofficial, may break at any time. Excessive use can trigger GitHub abuse detection. Use at your own risk.
TL;DR — Choose either supported runtime:
# Bun >= 1.4
bunx --bun ghc-proxy@latest start
# Node.js >= 24
npx ghc-proxy@latest startBefore you start, make sure you have:
- One supported JavaScript runtime:
- Bun >= 1.4:
winget install --id Oven-sh.Bunon Windows, or see the official installation guide - Node.js >= 24: install the latest LTS release from the official Node.js download page
- Bun >= 1.4:
- A GitHub Copilot subscription -- individual, business, or enterprise
-
Start the proxy with your chosen runtime:
# Bun bunx --bun ghc-proxy@latest start # Node.js npx ghc-proxy@latest start
-
On the first run, you will be guided through GitHub's device-code authentication flow. Follow the prompts to authorize the proxy.
-
Once authenticated, the proxy starts on
http://localhost:4141and is ready to accept requests.
That's it. Any tool that supports the OpenAI or Anthropic API can now point to http://localhost:4141.
The examples below use bunx --bun. If you chose Node.js, replace bunx --bun with npx; the published CLI and commands are the same.
Tip: If you set
--rate-limit, add--waitto queue requests instead of rejecting them with 429 when the cooldown has not elapsed yet. See Rate Limiting for details.
This is the most common use case. There are two ways to set it up:
bunx --bun ghc-proxy@latest start --claude-codeThis starts the proxy, opens an interactive model picker, and prints a ready-to-paste environment command. Run that command in another terminal to launch Claude Code with the correct configuration.
Create or edit ~/.claude/settings.json (this applies globally to all projects):
{
"env": {
"ANTHROPIC_BASE_URL": "http://localhost:4141",
"ANTHROPIC_AUTH_TOKEN": "dummy-token",
"ANTHROPIC_MODEL": "claude-opus-5",
"ANTHROPIC_DEFAULT_SONNET_MODEL": "claude-sonnet-5",
"ANTHROPIC_DEFAULT_HAIKU_MODEL": "claude-haiku-4.5",
"CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC": "1"
},
"permissions": {
"deny": ["WebSearch"]
}
}Then simply start the proxy and use Claude Code as usual:
bunx --bun ghc-proxy@latest startWhat each environment variable does:
| Variable | Purpose |
|---|---|
ANTHROPIC_BASE_URL |
Points Claude Code to the proxy instead of Anthropic's servers |
ANTHROPIC_AUTH_TOKEN |
Any non-empty string; the proxy handles real authentication |
ANTHROPIC_MODEL |
The model Claude Code uses for primary/Opus tasks |
ANTHROPIC_DEFAULT_SONNET_MODEL |
The model used for Sonnet-tier tasks |
ANTHROPIC_DEFAULT_HAIKU_MODEL |
The model used for Haiku-tier (fast/cheap) tasks |
CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC |
Disables telemetry and non-essential network traffic |
Tip: The model names above (e.g.
claude-opus-5) are mapped to actual Copilot models by the proxy. See Model Mapping below for details.
See the Claude Code settings docs for more options.
ghc-proxy uses a subcommand structure:
bunx --bun ghc-proxy@latest start # Start the proxy server
bunx --bun ghc-proxy@latest auth # Run GitHub auth flow without starting the server
bunx --bun ghc-proxy@latest check-usage # Show your Copilot usage/quota in the terminal
bunx --bun ghc-proxy@latest debug # Print diagnostic info (version, paths, token status)
bunx --bun ghc-proxy@latest selfcheck # Probe tokenizer chunks and Bun/Node runtime contracts in the packaged bundle| Option | Alias | Default | Description |
|---|---|---|---|
--port |
-p |
4141 |
Port to listen on (1..65535; malformed or zero values fail before startup) |
--verbose |
-v |
false |
Enable verbose logging |
--account-type |
-a |
individual |
individual, business, or enterprise |
--rate-limit |
-r |
-- | Minimum seconds between requests |
--wait |
-w |
false |
Queue requests instead of rejecting with 429 when --rate-limit cooldown has not elapsed (requires --rate-limit) |
--manual |
-- | false |
Manually approve each request |
--github-token |
-g |
-- | Use a GitHub token for this process only (normally obtained with auth); this flag does not persist it to config.json |
--claude-code |
-c |
false |
Generate a Claude Code launch command |
--show-token |
-- | false |
Display tokens on auth and refresh |
--dump-failed-payloads |
-D |
false |
Dump failed /responses payloads on upstream 400 errors for debugging. Can also be enabled with DUMP_FAILED_PAYLOADS=1. |
--proxy-env |
-- | false |
Use HTTP_PROXY/HTTPS_PROXY from env (Node.js only; Bun reads proxy env natively) |
--idle-timeout |
-- | 120 |
Bun server idle timeout in seconds (0 disables; Bun max is 255; streaming routes disable idle timeout automatically) |
--upstream-timeout |
-- | 1800 |
Upstream request timeout in seconds (0 disables). Enforced as a total-duration AbortSignal. Note both runtimes also apply their own ~300s idle timeout to fetch (Bun's built-in limit; Node's undici headersTimeout/bodyTimeout), which fires when no byte arrives for that long — a steadily streaming response is not capped by it, but a stalled one is rejected at ~300s and returned as a 504. |
--upstream-queue-concurrency |
-- | 10 |
Maximum concurrent Copilot upstream requests |
--upstream-queue-retries |
-- | 1 |
Maximum retries across capacity and approved pre-connection failures (0..2). Generation requests retry HTTP 429/529; effect-free requests retain the broader transient-status policy |
--upstream-recovery-budget |
-- | 60 |
Seconds available after the first approved retryable outcome or active-cooldown encounter for all later waits and pre-Response attempts (1..120) |
--upstream-queue-base-delay |
-- | 2 |
Base delay in seconds for upstream retry backoff when Retry-After is absent |
--upstream-queue-max-delay |
-- | 60 |
Maximum computed backoff in seconds; never shortens a valid Retry-After minimum |
--ghe-domain |
--ghe |
-- | GitHub Enterprise Cloud company domain (e.g. company.ghe.com). Required for GHE.com device login on first run; persisted automatically for later runs. |
If you want to throttle how often the proxy forwards requests:
# Enforce a 30-second cooldown between requests
bunx --bun ghc-proxy@latest start --rate-limit 30
# Same, but queue requests instead of returning 429
bunx --bun ghc-proxy@latest start --rate-limit 30 --wait
# Manually approve every request (useful for debugging)
bunx --bun ghc-proxy@latest start --manual--wait only takes effect when --rate-limit is also set. Without --rate-limit, there is no cooldown to wait on and --wait has no effect.
If you have a GitHub Business or Enterprise Copilot plan, pass --account-type:
bunx --bun ghc-proxy@latest start --account-type business
bunx --bun ghc-proxy@latest start --account-type enterpriseThis routes requests to the correct Copilot API endpoint for your plan. See the GitHub docs on network routing for details.
If your organization uses GitHub Enterprise Cloud (*.ghe.com), the standard GitHub device login URL differs from github.com. Pass your company's GHE domain on first auth:
bunx --bun ghc-proxy@latest start --account-type enterprise --ghe-domain company.ghe.comOr authenticate first, then start without the flag on subsequent runs:
# First run (authenticates and persists the domain)
bunx --bun ghc-proxy@latest auth --ghe-domain company.ghe.com
# Later runs (domain is read from persisted config)
bunx --bun ghc-proxy@latest start --account-type enterpriseThe proxy normalizes and persists the GHE domain automatically after a successful authentication, so you only need to pass --ghe-domain on the first run or when switching tenants.
Note:
--account-type enterprisealone is not sufficient for GHE.com login — the proxy needs the company domain to construct the correct device login URL (https://<company>.ghe.com/login/device). GHE.com support is scoped to*.ghe.comonly and does not apply to self-hosted GitHub Enterprise Server instances.
The proxy reads an optional JSON config file at:
~/.local/share/ghc-proxy/config.json
All fields are optional. The full schema:
| Field | Type | Default | Description |
|---|---|---|---|
githubToken |
string |
unset | Persisted GitHub token. The device-code flow (auth or first startup) writes it automatically; start --github-token is runtime-only and does not write this field |
modelRewrites |
{ from, to }[] |
[] |
Glob-pattern model substitution rules (see Model Rewrites) |
modelFallback |
object |
built-in family defaults | Override default model fallbacks (see Customizing Fallbacks) |
modelFallback.claudeOpus |
string |
claude-opus-5 |
Fallback for claude-opus-* models |
modelFallback.claudeSonnet |
string |
claude-sonnet-5 |
Fallback for claude-sonnet-* models |
modelFallback.claudeHaiku |
string |
claude-haiku-4.5 |
Fallback for claude-haiku-* models |
smallModel |
string |
unset | Target model for compact request routing (see Small-Model Routing) |
compactUseSmallModel |
boolean |
false |
Route compact/summarization requests to smallModel |
useFunctionApplyPatch |
boolean |
true |
Rewrite apply_patch custom tool as function tool on Responses path |
responsesApiAutoCompactInput |
boolean |
false |
Automatically trim Responses input to the latest compaction item |
responsesApiAutoContextManagement |
boolean |
false |
Automatically inject Responses context_management for selected models |
responsesApiContextManagementModels |
string[] |
[] |
Models eligible for auto-injected Responses context_management |
responsesApiParameterFilters |
{ models, params }[] |
[] |
Extra rules to strip request parameters on the Responses boundary; the built-in reasoning-model rule remains active unless replaced (see Responses Parameter Filters) |
responsesApiParameterFiltersReplaceDefault |
boolean |
false |
Disable the built-in reasoning-model default rule so only your responsesApiParameterFilters apply |
chatCompletionsUseMaxCompletionTokens |
string[] |
[] |
Extra model globs that rename Chat Completions max_tokens to max_completion_tokens; adds to the built-in gpt-5.4 / gpt-5.4-* rules |
responsesOfficialEmulator |
boolean |
false |
Enable local OpenAI-style Responses state emulation for previous_response_id, conversation, retrieve, input_items, delete, and input_tokens |
responsesOfficialEmulatorTtlSeconds |
number |
14400 |
In-memory TTL for locally emulated Responses state |
modelReasoningEfforts |
Record<string, string> |
{}; unlisted models use high |
Per-model reasoning effort defaults for Anthropic-to-Responses translation. Each value must be one of none, minimal, low, medium, high, xhigh, or max (ascending) |
upstreamQueueConcurrency |
number |
10 |
Maximum concurrent Copilot upstream requests |
upstreamQueueMaxRetries |
number |
1 |
Maximum retries across capacity and approved pre-connection failures (0..2) |
upstreamRecoveryBudgetSeconds |
number |
60 |
Shared recovery deadline after the first retryable outcome or active-cooldown encounter (1..120 seconds) |
overloadFallbacks |
Record<string, string> |
{} (disabled) |
Exact effective-model mappings for one opt-in fallback dispatch after terminal model 529 |
upstreamQueueBaseDelaySeconds |
number |
2 |
Base delay (seconds) for upstream retry backoff when Retry-After is absent |
upstreamQueueMaxDelaySeconds |
number |
60 |
Maximum computed backoff (seconds); does not clamp Retry-After |
gheDomain |
string |
unset | GitHub Enterprise Cloud company domain (persisted automatically after GHE.com auth) |
Example:
{
"modelRewrites": [
{ "from": "claude-haiku-*", "to": "gpt-4.1-mini" }
],
"modelFallback": {
"claudeOpus": "claude-opus-5",
"claudeSonnet": "claude-sonnet-5"
},
"smallModel": "gpt-4.1-mini",
"compactUseSmallModel": true,
"useFunctionApplyPatch": true,
"responsesApiAutoCompactInput": false,
"responsesApiAutoContextManagement": false,
"responsesApiContextManagementModels": ["gpt-5", "gpt-5-mini"],
"chatCompletionsUseMaxCompletionTokens": [],
"responsesOfficialEmulator": false,
"responsesOfficialEmulatorTtlSeconds": 14400,
"modelReasoningEfforts": {
"gpt-5": "high",
"gpt-5-mini": "medium"
},
"overloadFallbacks": {
"claude-opus-5": "claude-opus-4.8"
}
}Priority order for model fallbacks: environment variable > config.json > built-in default.
When Claude Code sends a request for a model like claude-sonnet-4.6, the proxy maps it to an actual model available on Copilot. The mapping logic works as follows:
- If the requested model ID is known to Copilot (e.g.
gpt-4.1,claude-sonnet-4.5), it is used as-is. - If the model starts with
claude-opus-,claude-sonnet-, orclaude-haiku-, it falls back to a configured model.
| Prefix | Default Fallback |
|---|---|
claude-opus-* |
claude-opus-5 |
claude-sonnet-* |
claude-sonnet-5 |
claude-haiku-* |
claude-haiku-4.5 |
You can override the defaults with environment variables:
MODEL_FALLBACK_CLAUDE_OPUS=claude-opus-5
MODEL_FALLBACK_CLAUDE_SONNET=claude-sonnet-5
MODEL_FALLBACK_CLAUDE_HAIKU=claude-haiku-4.5Or in the proxy's config file (~/.local/share/ghc-proxy/config.json):
{
"modelFallback": {
"claudeOpus": "claude-opus-5",
"claudeSonnet": "claude-sonnet-5",
"claudeHaiku": "claude-haiku-4.5"
}
}Note: Model fallbacks only apply to the chat completions translation path. The native Messages and Responses API strategies pass the model ID through to Copilot as-is.
For more general model substitution, use modelRewrites in the config file. Each rule maps a from pattern to a to model ID. The from field supports glob patterns with * wildcards, and the first matching rule wins.
{
"modelRewrites": [
{ "from": "claude-haiku-*", "to": "gpt-4.1-mini" },
{ "from": "gpt-5.4*", "to": "gpt-5.2" }
]
}Unlike model fallbacks (which only apply to the chat completions path), rewrites are applied uniformly to all three endpoints — /v1/messages, /v1/chat/completions, and /v1/responses. Target model names are normalized against Copilot's known model list using dash/dot equivalence (e.g. gpt-4.1 matches gpt-4-1).
Rewrites run before any other model policy — small-model routing and strategy selection all see the rewritten model.
overloadFallbacks is a separate, opt-in recovery policy. Each key and value is an exact advertised model ID after normal rewrite/compact resolution:
{
"overloadFallbacks": {
"claude-opus-5": "claude-opus-4.8"
}
}Fallback is considered only after a terminal source-model 529 or a pre-existing local cooldown for that source. It never runs for account 429, connection failures, timeouts, cancellation, validation failures, other statuses, or failures after an upstream Response exists. The target must be distinct, advertised, not locally cooled, and compatible with the request's endpoint, tools, parallel tools, streaming, vision, reasoning/thinking, and structured-output needs. The pipeline rebuilds target-dependent transforms and strategy selection from pristine input, dispatches once with no fresh retry allowance, and reports the actual served target in response model fields and the OVERLOAD_FALLBACK model trace.
Mappings are exact one-hop choices, not a traversed graph. Blank, same-model, and reciprocal two-node entries such as A -> B plus B -> A are ignored with a configuration warning. An unknown or incompatible runtime target preserves the source 529.
An upstream 429 establishes an account cooldown. A 529 is scoped to the final effective upstream model; if no effective model is known, it remains request-only. Eligible models can bypass cooled waiters while a global slot is free, but active slots and the maximum pending depth remain process-global limits.
The first attempt keeps the normal upstream timeout. The recovery budget starts at the first approved retryable outcome or the first encounter with an already-active cooldown, then covers every later cooldown/backoff wait, queue acquisition, same-model attempt until Response, and overload fallback. A valid integer-seconds, HTTP-date, or full-string fractional-seconds Retry-After is a strict lower bound and installs its full cooldown deadline. If that minimum cannot fit, the proxy skips the same-model retry instead of shortening it. Without a valid header, full jitter is sampled from zero through the smallest of exponential backoff, upstreamQueueMaxDelaySeconds, and remaining budget.
Generation requests additionally retry only measured pre-connection shapes: Bun ConnectionRefused, or Node ECONNREFUSED, ENOTFOUND, and EAI_AGAIN. Caller aborts, every timeout, TLS/configuration failures, resets, generic fetch errors, body failures, and stream failures are excluded. Once fetch() returns a Response, that attempt is committed: a later JSON/SSE/body failure is never replayed or sent to fallback.
Recovery logs use the existing request ID and structured fields such as event, retry count, status/connection class, effective model, scope, active/max slots, pending/max depth, queue wait, delay source/delay, elapsed/remaining budget, nextRetryAt, and decision. Public responses keep protocol-compatible payloads and only safe standard metadata such as Retry-After; no retry-progress SSE event or recovery payload extension is added. Per inbound request, the attempt ceiling is 1 + upstreamQueueMaxRetries + at most one configured fallback; outer SDK retries multiply that ceiling independently.
/v1/messages can optionally reroute specific low-value requests to a cheaper model:
smallModel: the model to reroute tocompactUseSmallModel: reroute recognized compact/summarization requests
The switch defaults to false. Routing is conservative:
- the target
smallModelmust exist in Copilot's model list - it must preserve the original model's declared endpoint support
- tool, thinking, and vision requests are not rerouted to a model that lacks the required capabilities
ghc-proxy sits between your tools and the GitHub Copilot API:
┌──────────────┐ ┌───────────┐ ┌───────────────────────┐
│ Claude Code │──────│ ghc-proxy │──────│ api.githubcopilot.com │
│ Cursor │ │ :4141 │ │ │
│ Any client │ │ │ │ │
└──────────────┘ └───────────┘ └───────────────────────┘
OpenAI or Translates GitHub Copilot
Anthropic between API
format formats
The proxy authenticates with GitHub using the device code OAuth flow (the same flow VS Code uses), then exchanges the GitHub token for a short-lived Copilot token that auto-refreshes.
When the Copilot token response includes endpoints.api, ghc-proxy now prefers that runtime API base automatically instead of relying only on the configured account type. This keeps enterprise/business routing aligned with the endpoint GitHub actually returned for the current token.
Incoming requests hit an Elysia server. chat/completions requests are validated, normalized into the shared planning pipeline, and then forwarded to Copilot. responses requests use a native Responses path with explicit compatibility policies. messages requests are routed per-model and can use native Anthropic passthrough, the Responses translation path, or the existing chat-completions fallback. The translator tracks exact vs lossy vs unsupported behavior explicitly; see the Messages Routing and Translation Guide and the Anthropic Translation Matrix for the current support surface.
The built-in, read-only Dashboard projects process health, model routing, behavior, and recent request lifecycle metadata without storing request or response content. See Dashboard Observability.
For Anthropic search_result blocks, an April 17, 2026 probe against claude-opus-4.6 on Copilot native /v1/messages accepted top-level search results and pure search-result tool outputs, but rejected top-level citations and mixed text/search-result tool output arrays. The native path sanitizes those observed rejection cases, while translated paths flatten search results to text; re-run the probe before treating that dated upstream result as universal.
ghc-proxy does not force every request through one protocol. The current routing rules are:
POST /v1/chat/completions: OpenAI Chat Completions -> shared planning pipeline -> Copilot/chat/completionsPOST /v1/responses: OpenAI Responses create -> native Responses handler -> Copilot/responsesPOST /v1/responses/input_tokens: Responses input-token counting passthrough by default, or local estimation in official emulator modeGET /v1/responses/:responseId: Responses retrieve passthrough by default, or local retrieval in official emulator modeGET /v1/responses/:responseId/input_items: Responses input-items passthrough by default, or local retrieval in official emulator modeDELETE /v1/responses/:responseId: Responses delete passthrough by default, or local deletion in official emulator modePOST /v1/messages: Anthropic Messages -> choose the best available upstream path for the selected model:- native Copilot
/v1/messageswhen supported - Anthropic -> Responses -> Anthropic translation when the model only supports
/responses - Anthropic -> Chat Completions -> Anthropic fallback otherwise
- native Copilot
This keeps the existing chat pipeline stable while allowing newer Copilot models to use the endpoint they actually expose.
OpenAI compatible:
| Method | Path | Description |
|---|---|---|
POST |
/v1/chat/completions |
Chat completions (streaming and non-streaming) |
POST |
/v1/responses |
Create a Responses API response |
POST |
/v1/responses/input_tokens |
Count Responses input tokens via upstream passthrough or the local official emulator |
GET |
/v1/responses/:responseId |
Retrieve one response via upstream passthrough or the local official emulator |
GET |
/v1/responses/:responseId/input_items |
Retrieve response input items via upstream passthrough or the local official emulator |
DELETE |
/v1/responses/:responseId |
Delete one response via upstream passthrough or the local official emulator |
GET |
/v1/models |
List available models |
POST |
/v1/embeddings |
Generate embeddings |
Anthropic compatible:
| Method | Path | Description |
|---|---|---|
POST |
/v1/messages |
Messages API with per-model routing across native Messages, Responses translation, or chat-completions fallback |
POST |
/v1/messages/count_tokens |
Token counting |
Utility:
| Method | Path | Description |
|---|---|---|
GET |
/health |
Liveness/readiness probe — returns { status, copilotToken, modelsLoaded, version } |
GET |
/usage |
Copilot quota / usage monitoring |
GET |
/token |
Inspect the current Copilot token |
Local Dashboard (read-only):
| Method | Path | Description |
|---|---|---|
GET |
/dashboard |
Dashboard application |
GET |
/dashboard/styles.css |
Dashboard stylesheet |
GET |
/dashboard/app.js |
Dashboard client script |
GET |
/dashboard/api/overview |
Process, authentication, quota, request, and queue summary |
GET |
/dashboard/api/models |
Upstream model metadata and effective proxy capabilities |
GET |
/dashboard/api/behavior |
Active routing, compatibility policies, strategies, and effect counters |
GET |
/dashboard/api/requests |
Active requests and the most recent 256 completed request summaries |
Dashboard routes are restricted to local access and return 403 when the peer, request host, or supplied Origin fails the loopback/same-origin checks. They are excluded from request history and access logging. See Dashboard Observability for the projection and security contract.
Note: The
/v1/prefix is optional for OpenAI-compatible endpoints (/chat/completions,/responsesand its resource routes,/models,/embeddings). Anthropic endpoints (/v1/messages,/v1/messages/count_tokens) require the/v1prefix. The utility and Dashboard endpoints are root-only and not exposed under/v1.
/v1/responses is designed to stay close to the OpenAI wire format while making Copilot limitations explicit:
- requests are validated before any mutation
- client-supplied
top_kis rejected with400on the OpenAI Chat Completions and Responses boundaries because neither official OpenAI schema defines it; clients that send it by mistake receive an explicit error instead of a silent drop. Anthropic Messagestop_kremains supported and is preserved when the proxy translates that request internally for Copilot - common official request fields such as
conversation,previous_response_id,max_tool_calls,truncation,user,prompt, andtextare now modeled explicitly instead of relying on loose passthrough alone - official
text.formatoptions are modeled explicitly, includingtext,json_object, andjson_schema - an opt-in
responsesOfficialEmulatormode adds in-memory OpenAI-style state forprevious_response_id,conversation,GET /responses/{id},GET /responses/{id}/input_items,DELETE /responses/{id}, andPOST /responses/input_tokens - emulator state is memory-only and expires after
responsesOfficialEmulatorTtlSeconds(default14400, or 4 hours) background: trueis rejected explicitly while emulator mode is enabledcustomapply_patchcan be rewritten as a function tool whenuseFunctionApplyPatchis enabled- automatic Responses
context_managementinjection is disabled by default and only applies whenresponsesApiAutoContextManagementistrueand the model matchesresponsesApiContextManagementModels - automatic trimming of Responses
inputto the latestcompactionitem is disabled by default and only applies whenresponsesApiAutoCompactInputistrue - reasoning defaults for Anthropic -> Responses translation can be tuned with
modelReasoningEfforts - request parameters that a model rejects (e.g.
temperature/top_pon reasoning models) are stripped on the Responses boundary rather than leaked upstream as a400; see Responses Parameter Filters - built-in web search (
web_search,web_search_preview, and their dated variants) is forwarded to Copilot rather than blocked; every/responsesmodel reached by the August 4, 2026 acceptance sweep accepted the tool, while functional search execution was verified ongpt-5.6-solandgpt-5.6-terra, see docs/research/responses-web-search.md - external image URLs on the Responses path fail explicitly with
400; usefile_idor data URL image input instead - official
input_fileanditem_referenceinput items are modeled explicitly and validated, but the verified Copilot GPT Responses boundary is stateless: it rejectsstore: trueand cannot resolve returned item IDs on later requests. The proxy deliberately applies a proxy-widestore: falsepolicy, removes allitem_referenceitems before dispatch, and removesfunction_call_outputitems whosecall_idhas no matchingfunction_callin the same input array. Without the optional emulator, a caller that requested storage still receives a successful stateless response; retrieve/delete/continuation semantics are available only from the local emulator
Example opt-in configuration for these two Responses-specific policies:
{
"responsesApiAutoContextManagement": true,
"responsesApiContextManagementModels": ["gpt-5"],
"responsesApiAutoCompactInput": true,
"responsesOfficialEmulator": true,
"responsesOfficialEmulatorTtlSeconds": 14400
}See Responses Upstream Notes for detailed upstream compatibility observations from live testing.
Some Copilot models reject request parameters that the OpenAI wire format allows. The clearest case: reasoning models (the gpt-5 family, o-series, codex) reject sampling parameters and answer POST /responses with 400 Unsupported parameter: 'temperature' is not supported with this model. Since the client cannot always be changed, the proxy strips the offending parameters on the Responses boundary instead of leaking the incompatibility outward.
This is expressed as a small rule engine that runs on both the native /v1/responses path and the /v1/messages → Responses translation path:
- Built-in default rule: any model that advertises
reasoning_efforthastemperaturestripped. It also hastop_pstripped except for*-codex/*-codex-*models, which are exempt because the July 26, 2026 probe found the tested Codex model acceptedtop_pwhile its reasoning-model siblings rejected it. This exemption narrows only the built-in rule; an operator rule can still striptop_p. responsesApiParameterFilters: add your own rules. Each rule is{ "models": [glob, ...], "params": [name, ...] }; every rule whosemodelsglob matches the resolved model contributes itsparams. Rules are added to the default (the union of parameters is stripped). Model globs use the same*wildcard asmodelRewrites.responsesApiParameterFiltersReplaceDefault: set totrueto disable the built-in reasoning-model rule, so only yourresponsesApiParameterFiltersapply — use this to fully overwrite the default behavior.
Stripped parameters are removed entirely (never sent as null), because upstream rejects the mere presence of the key.
{
"responsesApiParameterFilters": [
{ "models": ["gpt-5*", "o1*"], "params": ["temperature", "top_p"] },
{ "models": ["some-model"], "params": ["service_tier"] }
],
"responsesApiParameterFiltersReplaceDefault": false
}Pre-built images are available on GHCR:
docker pull ghcr.io/wxxb789/ghc-proxy
docker volume create ghc-proxy-data
docker run --rm -p 127.0.0.1:4141:4141 \
-v ghc-proxy-data:/home/bun/.local/share/ghc-proxy \
ghcr.io/wxxb789/ghc-proxyOr build locally:
docker build -t ghc-proxy .
docker volume create ghc-proxy-data
docker run --rm -p 127.0.0.1:4141:4141 \
-v ghc-proxy-data:/home/bun/.local/share/ghc-proxy \
ghc-proxyAuthentication and settings are persisted in the ghc-proxy-data volume so they survive container restarts. The proxy does not provide API authentication. Keep the port bound to loopback as shown; any non-loopback deployment needs an authenticated TLS reverse proxy or a firewall that restricts access.
Run the device-code authentication flow once against the same volume:
docker run --rm -it \
-v ghc-proxy-data:/home/bun/.local/share/ghc-proxy \
ghcr.io/wxxb789/ghc-proxy authThe legacy --auth container argument remains supported, but auth is the standard CLI subcommand:
docker run --rm -it \
-v ghc-proxy-data:/home/bun/.local/share/ghc-proxy \
ghcr.io/wxxb789/ghc-proxy --authYou can also pass a GitHub token via GH_TOKEN. The container entrypoint forwards a non-empty value only when starting the proxy, as start --github-token:
docker run --rm -p 127.0.0.1:4141:4141 \
-v ghc-proxy-data:/home/bun/.local/share/ghc-proxy \
-e GH_TOKEN=your_token \
ghcr.io/wxxb789/ghc-proxyDocker Compose:
services:
ghc-proxy:
image: ghcr.io/wxxb789/ghc-proxy
ports:
- '127.0.0.1:4141:4141'
volumes:
- ghc-proxy-data:/home/bun/.local/share/ghc-proxy
environment:
- GH_TOKEN=your_token_here
restart: unless-stopped
volumes:
ghc-proxy-data:Repository development uses Bun >= 1.4 even if you run the published package with Node.js.
git clone https://github.com/wxxb789/ghc-proxy.git
cd ghc-proxy
bun install
bun run dev # Start with --watch
# Or use the production-style source command:
bun run startbun install # Install dependencies
bun run dev # Start with --watch
bun run start # Start without --watch
bun run build # Build with tsdown
bun run lint # ESLint
bun run typecheck # tsc --noEmit
bun test # Run tests
bun run matrix:live # Real Copilot upstream compatibility matrix
bun run matrix:live --vision-only --all-responses-models --json
bun run matrix:live --stateful-only --json --model=gpt-5.2-codexNote:
bun run matrix:liveuses your configured GitHub/Copilot credentials and spends real upstream requests. Use it when you want end-to-end verification against the current Copilot service, not for every local edit.Useful flags:
--json: emit machine-readable JSON only--vision-only: run just the Responses image probes--stateful-only: run follow-up/resource probes such asprevious_response_id,input_tokens, andinput_items--all-responses-models: scan every model that advertises/responses--model=<id>: pin the Responses scan to one specific model
Answers whether a tool works, not merely whether upstream returns 200 when you mention it. For each (model × tool) it declares the tool, then — unless --accept-only — sends a prompt that cannot be answered without it and looks in the response for proof it ran.
Verdicts: supported (the tool ran or was called), inert (accepted but never invoked), unsupported (upstream rejected it), unmeasured (capacity/gateway fault — no verdict, re-run before publishing).
bun scripts/probes/tool-support.ts # both boundaries
bun scripts/probes/tool-support.ts --json # JSON snapshot to stdout
bun scripts/probes/tool-support.ts --model=claude-opus-5 # single model
bun scripts/probes/tool-support.ts --boundary=responses # or: messages
bun scripts/probes/tool-support.ts --accept-only # skip the functional pass (half the quota)
bun scripts/probes/tool-support.ts --names # also probe client tool NAMES (WebSearch, shell, ...)Latest results: docs/research/builtin-tool-support.md.
The JSON output is designed for weekly diffing — generatedAt is the only volatile field:
# Compare two weekly snapshots
diff <(jq -S 'del(.generatedAt)' week1.json) <(jq -S 'del(.generatedAt)' week2.json)