One local control plane
A single localhost gateway becomes the stable front door for supported backends. Clients do not need to know which adapter handles a particular model.
Infera gives local AI a polished front door. Chat with supported model backends, route compatible requests through one local gateway, and keep the runtime on your own machine.
Infera is designed so a curious developer can understand what it does before running it. Every visual layer maps back to an actual capability.
A single localhost gateway becomes the stable front door for supported backends. Clients do not need to know which adapter handles a particular model.
Responses can arrive incrementally instead of waiting for a finished answer.
SSEModel prefixes and backend rules determine where a request goes.
adaptersFamiliar API shapes reduce the amount of client-side rework.
OpenAI + AnthropicThe default gateway binds to localhost so the control plane stays on your machine.
127.0.0.1The architecture treats model backends as adapters. That separation makes the public interface more stable while backend implementations can evolve independently.
Web-backed Gemini integration exposed behind the local gateway boundary.
backend adapterQwen-compatible backend surface connected through the routing layer.
backend adapterA compatible client-facing contract can remain stable while the backend changes underneath.
compatibility layerInfera exposes familiar request shapes for developers who already work with model APIs, while keeping the service addressable on localhost.
Representative routes exposed by the gateway.
GET /v1/modelsdiscover modelsPOST /v1/chat/completionsOpenAI-style chatPOST /v1/messagesAnthropic-style messagesPOST /v1/responsesResponses surfacePOST /v1/images/generationsimage generation surfacePoint an OpenAI-compatible SDK at your local gateway and let Infera handle backend routing.
No magic box. The system is easiest to understand as a sequence of explicit boundaries.
The website explains the architecture, but the repository remains the source of truth for the implementation, licenses, security notes, and setup.
Read the repository before installation. Infera is designed to be understandable instead of opaque.
GitHubThe gateway's default entry point is localhost. The public page is an interface; it does not become your inference server.
localhostInfera retains explicit attribution to Sophomoresty and the original gemini-web2api project where applicable.
upstreamExplore the source, inspect the architecture, or open the local workspace experience.
Infera is meant to sit between a polished developer experience and the real machinery of local model access. The point is not to hide complexity; it is to organize it.
Each model backend has its own client assumptions, routes, streaming details, and operational quirks.
Clients talk to one local surface. Infera handles discovery, routing, compatibility, and streaming boundaries.
Choose the path that matches how much control you want. The source remains available either way.
Download the launcher and run it locally.
Download .CMDUse the shell installer from the repository.
Open installerClone or browse the repository before running anything.
View sourceBase URL: http://127.0.0.1:8081/v1. Use the gateway as the stable local boundary and let backend adapters handle provider-specific behavior.
This is the browser-side experience. Start the local gateway on your machine to make the real connection.