Provider integration

A governed gateway for every agent call

Kimss is the control plane. Azure AI Foundry is one supported data plane—alongside OpenAI-compatible BYO. Workspace identity, tenant routing, usage attribution, and governed requests sit in front of whichever runtime you connect.

Last updated: July 23, 2026

Execution plane and control plane

Foundry runs agents and models; Kimss routes authenticated workspace traffic, meters governed requests, and enforces tenant boundaries on every call through the universal gateway.

Clients no longer need separate Foundry credentials per microservice. Kimss maps X-Kimss-Key or bearer tokens plus optional X-Workspace-ID to the correct Foundry project or shared shard.

Route by workspace

Map users and API keys to Foundry projects without blending tenants.

Meter consistently

Meter model and agent traffic as governed requests with agent-level attribution.

Enforce boundaries

Keep workspace identity attached instead of one shared Azure resource key.

One preferred developer surface

POST /v1/agents/run and POST /v1/models/completions form the supported contract documented at /docs/api_docs—authenticated with Kimss keys and workspace headers.

Streaming, tool calls, and vector retrieval flow through the same gateway. Legacy assistant endpoints remain for migration periods; new code should target /v1.

Optional Azure API Management can front AI traffic when telemetry and policies are verified end to end. Direct Foundry routing remains the safe default documented in /docs/architecture.

Operations at scale

Platform teams monitor pool exhaustion, rotate workspace keys, and align Foundry project capacity with credit policy—using Kimss admin surfaces plus Azure Monitor.

When enabling APIM GatewayLogs, pair gateway diagnostics with Kimss attribution for investigations. Do not enable production APIM gateway mode until required E2E tests pass in your environment.

Start integrating

Create a workspace, run the Python SDK quickstart, and migrate legacy assistant traffic using the compatibility guide when you expand beyond pilots.

Related: /enterprise-ai-control-plane, /azure-ai-foundry-governance, /ai-credit-governance.

Implementation patterns for the agent gateway

Teams succeed with the agent gateway when they treat Kimss as the integration boundary: applications never hold provider secrets, every call includes workspace context, and operators review governed-request trends before expanding model access.

Start in a non-production workspace. Wire gateway-authenticated calls against POST /v1/agents/run or POST /v1/models/completions using X-Kimss-Key or a bearer token. Validate streaming, tool invocation, and error paths your production clients rely on.

Document which Entra groups map to which workspace roles. Align workspace routing tables with monthly governed-request allowances so finance sees predictable units rather than surprise token spikes on the Azure invoice.

Publish an internal integration checklist: required headers, workspace identifiers, approved models, and escalation paths when governed-request allowances approach exhaustion.

Use /docs/architecture to confirm the APIM byo-proxy path for customer inference. Foundry ai-proxy is not the customer default.

When shared Foundry infrastructure spans multiple internal products, give each product its own API key or sub-workspace budget so attribution stays legible in usage aggregates and billing ledgers.

Common mistakes when rolling out the agent gateway

The costliest errors are shared Foundry keys in microservices, skipping workspace headers on multi-tenant keys, and migrating user-facing flows before server-side credit enforcement is tested.

Embedding one project key in every service bypasses Kimss RBAC and makes revocation a company-wide fire drill. Issue workspace-scoped keys per service or per environment instead.

Assuming legacy /assistant_* behavior matches /v1 governance causes silent gaps in metering or identity. Inventory clients with /assistants-to-v1-migration and retire legacy paths deliberately.

Treating governed-request limits as cosmetic reporting rather than enforced allowances invites overrun. Configure exhaustion policies in staging and confirm blocked requests behave as product management expects.

Publishing internal runbooks that reference production Swagger instead of /docs/api_docs creates integration drift. The public API reference is the supported contract for external integrators.

Skipping staging verification for streaming and tool calls leads to production surprises. Exercise the same client libraries and timeouts you expect under peak load.

Next steps for the agent gateway

Create a workspace, read /why-kimss for positioning, follow /python-sdk-mcp-quickstart for code, and engage /enterprise when contractual isolation, capacity, or onboarding differ from self-serve plans.

Self-serve teams typically progress: signup, first agent run via SDK, governed-request allowance configuration, Entra SSO for Studio users, then wider rollout to internal consumers or customer tenants.

For the agent gateway, schedule a monthly review of usage aggregates, ledger entries, and agent inventory. Remove unused keys, archive obsolete agents, and adjust group budgets after major launches.

Customer-facing ISVs should pair Kimss workspace design with /multi-tenant-ai-security and /ai-rbac-and-identity so each end customer receives isolated agents, files, and usage rows.

Track product changes at /changelog and deeper narratives at /insights so your platform team does not miss SDK or API shifts that affect deployed clients.

Documentation and honest scope

Kimss documents the supported integration surface at /docs/api_docs and system design at /docs/architecture—avoid assuming every internal admin route is available in the public SDK or MCP server.

Platform engineers should bookmark /docs/api_docs as the contract for external integrators. When product management requests a feature, verify whether it exists on /v1, requires an admin API, or needs net-new development before committing customer timelines.

governed requests, Entra SSO, workspace RBAC, and PostgreSQL isolation are first-class product capabilities—not marketing adjectives. Validate them in your tenant with test workspaces and realistic agent workloads rather than slide-deck assumptions.

Optional Azure API Management integration remains documented as an advanced path. Production enablement should follow your organization's verification checklist for gateway telemetry and routing parity with direct Foundry execution.

When questions fall outside public documentation, enterprise customers can reach Kimss via /enterprise. Self-serve builders can use in-product support after signup.

Verify before you scale

Treat Kimss as production infrastructure: validate identity, governed-request caps, and routing in a staging workspace, read /docs/api_docs for the supported contract, and expand pools only after usage patterns are understood.

Platform teams should run monthly reviews of workspace keys, agent inventory, and ledger entries. Remove unused credentials, archive obsolete agents, and align governed-request allowances with teams that actually ship. Pair Kimss attribution with Azure Cost Management for infrastructure truth—Kimss meters governed requests on the control plane; Azure (or your provider) still bills underlying model consumption.

When you need help beyond public documentation, self-serve builders use in-product support after signup; enterprise buyers start at /enterprise for onboarding, capacity, and contractual questions. Product changes publish at /changelog so integrators can track SDK and API shifts over time.