Release Date: April 2026
Release 0.7 is a major release focused on completing the transition to OpenAI API conformance, introducing comprehensive observability metrics, and significant API cleanup. This release removes the fine-tuning API, completes the FastAPI router migration, removes legacy providers (TGI, HuggingFace), renames core concepts for clarity (agents to responses, knowledge_search to file_search, meta-reference to builtin), and adds structured logging via structlog. New providers include Infinispan vector-io and inline Docling for PDF parsing.
Note: Each change is described in detail with migration instructions in the sections below.
Hard Breaking Changes (action required before upgrading):
| Change | Migration | PR |
|---|---|---|
| Fine-tuning API removed | Remove all /post-training and /fine-tuning API usage |
#5104 |
meta-reference providers renamed to builtin |
Replace inline::meta-reference with inline::builtin in configs |
#5131 |
knowledge_search renamed to file_search |
Replace knowledge_search with file_search in tool names and API calls |
#5186 |
| Agents API renamed to Responses API | Replace /agents endpoints with /responses |
#5195 |
tool_groups removed from public API |
Remove tool_groups registration; providers auto-register tools |
#4997 |
| TGI and HuggingFace inference providers removed | Switch to remote::vllm, remote::ollama, or other providers |
#5333 |
register/unregister model endpoints removed |
Use standard CRUD endpoints for model management | #5341 |
@webmethod decorator removed |
All APIs now use FastAPI routers exclusively | #5248 |
rag-runtime provider renamed to file-search |
Replace inline::rag-runtime with inline::file-search in configs |
#5187 |
Duplicate dataset_id parameter removed from append-rows |
Remove dataset_id from request body (use path parameter only) |
#4849 |
/files/{file_id} GET response format unified |
Update clients expecting provider-specific response shapes | #5154 |
| OpenAI API schema transforms added | Review response schemas for new conformance fields | #5166 |
starter-gpu distribution removed |
Use starter distribution with remote inference providers |
#5279 |
Behavior Changes (no code changes required, but be aware):
| Change | Note | PR |
|---|---|---|
sentence_transformers trust_remote_code defaults to False |
Set trust_remote_code: true in config if using custom models |
#4602 |
| Inline neural rerank for RAG | Reranking now available as built-in capability | #4877 |
| Logging migrated to structlog | Log output format changed to structured key-value pairs | #5215 |
These changes take effect immediately and require updates before upgrading.
1. Fine-Tuning API Removed (#5104)
Contributed by Sebastien Han (Red Hat) — @leseb
Impact: All users of the post-training/fine-tuning API
The entire fine-tuning API has been removed from OGX. This includes all /post-training endpoints and related provider implementations.
Migration: Remove all fine-tuning API calls. Use external fine-tuning services directly.
2. meta-reference Providers Renamed to builtin (#5131)
Contributed by Sebastien Han (Red Hat) — @leseb
Impact: All users with meta-reference in their distribution configs
Before:
provider_type: inline::meta-referenceAfter:
provider_type: inline::builtinMigration: Search and replace inline::meta-reference with inline::builtin in all config files.
grep -r "meta-reference" your-config-directory/3. knowledge_search Renamed to file_search (#5186)
Contributed by Sebastien Han (Red Hat) — @leseb
Impact: All users referencing knowledge_search tool name
Before:
tools=[{"type": "knowledge_search", ...}]
After:
tools=[{"type": "file_search", ...}]
Migration:
grep -r "knowledge_search" your-project/4. Agents API Renamed to Responses API (#5195)
Contributed by Sebastien Han (Red Hat) — @leseb
Impact: All users of the agents API endpoints
The agents API has been renamed to the Responses API to align with OpenAI's naming convention. All /agents endpoints are now served under /responses.
Migration: Update API endpoint paths from /agents/* to /responses/*.
5. tool_groups Removed from Public API (#4997)
Contributed by Sebastien Han (Red Hat) — @leseb
Impact: Users who manually register tool groups
Tool groups are now auto-registered from provider specs. The public API for registering/unregistering tool groups has been removed.
Migration: Remove any tool_groups registration calls. Ensure your provider specs include toolgroup_id where needed.
6. TGI and HuggingFace Inference Providers Removed (#5333)
Contributed by Sebastien Han (Red Hat) — @leseb
Impact: Users relying on remote::tgi or remote::huggingface providers
Migration: Switch to remote::vllm, remote::ollama, or another supported inference provider.
7. Deprecated register/unregister Model Endpoints Removed (#5341)
Contributed by Sebastien Han (Red Hat) — @leseb
Impact: Users calling deprecated model registration endpoints
Migration: Use standard model management endpoints.
8. @webmethod Decorator Removed — FastAPI Router Migration Complete (#5248)
Contributed by Sebastien Han (Red Hat) — @leseb
Impact: Custom provider authors who used @webmethod
All APIs now use FastAPI routers exclusively. The legacy @webmethod decorator has been removed.
Migration: Convert any custom endpoints to use FastAPI router decorators.
9. rag-runtime Provider Renamed to file-search (#5187)
Contributed by Sebastien Han (Red Hat) — @leseb
Impact: Users with inline::rag-runtime or builtin::rag in configs
Before:
provider_type: inline::rag-runtime
toolgroup_id: builtin::ragAfter:
provider_type: inline::file-search
toolgroup_id: builtin::file-search10. Duplicate dataset_id Removed from Append-Rows (#4849)
Contributed by Eoin Fennessy (Red Hat) — @eoinfennessy
Impact: API clients sending dataset_id in both the URL path and request body
Migration: Remove dataset_id from the request body; it is now taken exclusively from the URL path parameter.
11. /files/{file_id} GET Response Unified (#5154)
Contributed by @r3v5
Impact: Clients parsing file metadata responses
The GET endpoint now returns a consistent response shape across all providers, eliminating provider-specific differences.
12. OpenAI API Schema Transforms (#5166)
Contributed by Nathan Weinberg (Red Hat) — @nathan-weinberg
Impact: Clients relying on the exact API response schema
Schema transforms and new types have been added to improve OpenAI API conformance. Response shapes may differ from previous versions.
13. starter-gpu Distribution Removed (#5279)
Contributed by Sebastien Han (Red Hat) — @leseb
Impact: Users of the starter-gpu distribution
Migration: Use the starter distribution with remote inference providers instead.
sentence_transformers trust_remote_code Now Defaults to False (#4602)
Contributed by Derek Higgins (Red Hat) — @derekhiggins
For security, trust_remote_code now defaults to False for sentence_transformers models. If you use custom models that require remote code execution, set trust_remote_code: true in your provider config.
Structured Logging via structlog (#5215)
Contributed by Sebastien Han (Red Hat) — @leseb
All logging has been migrated to structlog with structured key-value output. Log parsing tools that rely on the previous format may need to be updated.
- Reasoning output in Responses API — Models can now return reasoning/thinking traces as part of response output (#5206 by @robinnarsinghranabhat)
- Reasoning as valid conversation item — Reasoning traces can be included in conversation history (#5392)
- Reasoning output types in OpenAI spec — Added
reasoningoutput type to the Responses API spec (#5357)
- API-level request metrics — Track request counts, latency, and error rates at the API layer (#5201 by @gyliu513)
- Inference metrics — Token throughput, latency, and model-level metrics (#5320 by @gyliu513)
- Vector IO metrics — Performance metrics for vector store operations (#5096 by @gyliu513)
- Parameter usage metrics for Responses API — Track parameter usage patterns (#5255 by @gyliu513)
Inline Neural Rerank for RAG (#4877)
Contributed by @r3v5
RAG pipelines can now use built-in neural reranking without external services, improving search quality with cross-encoder models.
Inline Docling Provider for PDF Parsing (#5049)
Contributed by @alinaryan
Structure-aware PDF parsing using Docling for high-quality document ingestion into vector stores.
Background Response Cancellation (#5268)
Contributed by Charlie Doern (Red Hat) — @cdoern
Added a cancel endpoint for background responses, allowing clients to abort long-running response generation.
Connector API Promoted to v1beta (#5129)
Contributed by Sebastien Han (Red Hat) — @leseb
The Connector API for MCP server management has been promoted from v1alpha to v1beta, signaling increased API stability.
stream_options Parameter Support (#4815)
Contributed by @gyliu513
Added support for the stream_options parameter in chat completions, enabling include_usage in streaming responses for OpenAI conformance.
Forward Headers for Inference Passthrough (#5134)
Contributed by @skamenan7
The inference passthrough provider now supports forwarding custom headers to upstream backends.
Form-Encoded Content Type for Responses API (#5193)
Contributed by @r3v5
The Responses API now accepts application/x-www-form-urlencoded content type in addition to JSON.
PGVector Filter Support (#5111)
Contributed by @franciscojavierarceo
PGVector now supports metadata filters for vector search queries, and f-string usage in table names has been replaced for safety.
Configurable asyncpg Connection Pools (#5160)
Contributed by @iamemilio
PostgreSQL connection pool settings (min/max connections, timeouts) are now configurable via provider config.
Contributed by Sebastien Han (Red Hat) — @leseb
A provider compatibility matrix for the Responses API and provider version tracking have been added to help users understand which features each provider supports.
Responses API Test Coverage Analyzer (#5101)
Conformance annotations and a test coverage analyzer for the Responses API have been added to track OpenAI specification coverage.
Infinispan Vector-IO Provider (#4839)
Contributed by @rigazilla
New vector store provider using Infinispan for distributed, high-performance vector storage.
- WatsonX: LiteLLM replaced with OpenAI mixin — Cleaner, more maintainable WatsonX provider (#5133 by @c99cd1a93)
file_searchdecoupled from legacyknowledge_searchtool_groups (#5175 by @leseb)- Large files split into focused modules (#5281, #5299 by @leseb)
- Tools API converted to FastAPI router (#5246 by @leseb)
- Unused
LiteLLMOpenAIMixinremoved (#5159)
- Lazy-load numpy, faiss, sqlite_vec in vector_io providers to reduce startup memory (#5118)
- Lazy-load torch and transformers in prompt_guard (#5117)
- Lazy-load torch in embedding_mixin to reduce startup memory (#5116)
- Lazy-load braintrust autoevals to reduce idle memory (~63MB) (#5078)
- Path traversal and header injection defenses (#5086 by @rhdedgar)
- CVE-2026-33236: Bump nltk to 3.9.4 (#5259 by @eoinfennessy)
- CVE-2026-30922: Bump pyasn1 to 0.6.3 (#5207 by @eoinfennessy)
- CVE-2026-32597: Bump pyjwt to 2.12.0 (#5127 by @eoinfennessy)
- Fix provider_data_var context leak (#5227 by @jaideepr97)
- Prevent OTel context leak in fire-and-forget background tasks (#5168 by @iamemilio)
- Disable asyncpg OTel auto-instrumentation to prevent duplicate DB spans (#5158)
- Fix asyncio event loop mismatch via operation deferral (#5130 by @derekhiggins)
- Improve chat completions OpenAI conformance (#5108 by @cdoern)
- Multi-worker cache synchronization for vector stores (#5076 by @elinacse)
- Replace blocking
requestscalls with asynchttpxin WatsonX (#5280 by @gyliu513) and remote providers (#5162 by @gyliu513) - Use SDK-native model names for Vertex AI (#5169 by @major)
- Fix vLLM health() and rerank() TLS and auth credentials (#5340 by @gyliu513)
- Fix vLLM rerank() provider-data-aware API key lookup (#5374)
- Fix require_approval field check in ApprovalFilter (#5288 by @jaideepr97)
- Gate conversation sync on store flag to prevent data leak when store=false (#5305 by @jaideepr97)
- Fix assistant message rewriting in _separate_tool_calls (#5303 by @jaideepr97)
- Handle asyncio.CancelledError in metrics try/except blocks (#5336 by @gyliu513)
- Race condition fix in background response cancel (#5363 by @leseb)
- Allow multi-worker server with dual-stack IPv6 support (#5284 by @derekhiggins)
- Auto-expand provider dependencies for
--providersin stack CLI (#4654 by @gyliu513) - Make InmemoryKVStore.delete consistent with other backends (#5289 by @gyliu513)
- Convert Path to str in _build_ssl_context() for httpx compatibility (#5380 by @gyliu513)
- Surface tiktoken encoding check at provider startup (#5401 by @Bobbins228)
- Pre-cache tiktoken cl100k_base encoding at image build time (#5391 by @Bobbins228)
- Fix milvus hybrid ranker usage (#5312 by @jakub-walaszczyk)
- Remove duplicate decode (#5177)
- Optimize connector listing (#5164)
- Remove references to defunct inline::builtin inference provider (#5174 by @leseb)
- Rewrite README and docs to lead with OpenAI API compatibility (#5323)
- Add AGENTS.md with guidelines for AI coding agents (#5211)
- Add architecture documentation and module-level READMEs (#5213)
- Blog post on Open Responses compliance and OpenAI compatibility (#5232)
- Blog post on OGX observability (#5387)
- Agentic flows tutorial blog post (#5035)
- Blog post about Responses API (#5196)
- Docling provider setup and usage docs (#5329)
- Multi-tenant isolation example for conversations and responses (#5176)
- Update stale documentation to reflect current architecture (#5393)
- Mintlify-inspired documentation UI improvements (#5405)
- Add docstrings to public classes and functions (#5267)
- Update README badges with logos, conformance score, and DeepWiki (#5389)
- Auto-record integration tests on PRs with multi-provider support (#5123)
- Add Bedrock to responses CI suite with recordings (#5254)
- Add WatsonX Responses API integration test recordings (#5120)
- Test Responses API against Azure AI Foundry (#5107)
- Add GCP Workload Identity Federation for Vertex AI recording workflow (#5276)
- Cache HuggingFace models and datasets for offline replay tests (#5382)
- Add conventional-pre-commit for commit validation (#5251)
- Add markdownlint and actionlint pre-commit hooks (#5271, #5285)
- Add mypy pre-commit hook enforcement (#5269)
- Replace Mergify queue with GitHub merge queue (#5383)
- Test last 3 release branches in scheduled CI (#5277)
- Remove docker mode from integration test matrix (#5311)
These hard breaking changes require updates before you can run 0.7:
-
Check for fine-tuning API usage:
grep -r "post-training\|fine.tuning\|fine_tuning" your-project/Remove all fine-tuning API calls.
-
Update provider names in configs:
grep -r "meta-reference" your-config-directory/ grep -r "rag-runtime" your-config-directory/ grep -r "starter-gpu" your-config-directory/
inline::meta-reference->inline::builtininline::rag-runtime->inline::file-searchbuiltin::rag->builtin::file-searchstarter-gpu->starter
-
Update tool names:
grep -r "knowledge_search" your-project/Replace with
file_search. -
Update API endpoints:
grep -r "/agents" your-project/Replace agents API calls with responses API equivalents.
-
Remove tool_groups registration:
grep -r "tool_groups\|register_tool" your-project/Tool groups are now auto-registered from provider specs.
-
Check for removed providers:
grep -r "remote::tgi\|remote::huggingface" your-config-directory/Switch to
remote::vllm,remote::ollama, or another supported provider. -
Check for deprecated model endpoints:
grep -r "register_model\|unregister_model" your-project/
- Review log output if you have log parsing tools, as logging now uses structured key-value format via structlog.
- If using
sentence_transformerswith custom models requiring remote code, addtrust_remote_code: trueto your provider config.