LOCAL-FIRST · OPEN SOURCE · UNIFIED AI LAYER

One interface.
Every model.

Infera gives local AI a polished front door. Chat with supported model backends, route compatible requests through one local gateway, and keep the runtime on your own machine.

Explore the system
Your browser is the interface. Your machine is the runtime. GitHub Pages never hosts your inference.
Gateway127.0.0.1:8081
Geminiadapter backend
Qwenadapter backend
OpenAI-compatibledeveloper surface
Infera Coreroute · stream · unify
Gateway127.0.0.1:8081local entry point
InterfaceBrowser-firstchat + API tooling
ProtocolsOpenAI · Anthropiccompatibility layers
RuntimeYour machineno hosted inference on the site
01 · Product

Beautiful on the surface.
Transparent underneath.

Infera is designed so a curious developer can understand what it does before running it. Every visual layer maps back to an actual capability.

◈

One local control plane

A single localhost gateway becomes the stable front door for supported backends. Clients do not need to know which adapter handles a particular model.

Clientbrowser · SDK · curl
→
Inferaroute + normalize
→
AdapterGemini / Qwen / etc.
→
Responsestream / JSON
POST http://127.0.0.1:8081/v1/chat/completions { "model": "your-model", "messages": [{"role":"user","content":"Hello"}], "stream": true }
⌁

Streaming

Responses can arrive incrementally instead of waiting for a finished answer.

SSE
⌘

Routing

Model prefixes and backend rules determine where a request goes.

adapters
▣

Compatibility

Familiar API shapes reduce the amount of client-side rework.

OpenAI + Anthropic
◇

Local-first

The default gateway binds to localhost so the control plane stays on your machine.

127.0.0.1
02 · Models

Different backends.
One developer experience.

The architecture treats model backends as adapters. That separation makes the public interface more stable while backend implementations can evolve independently.

✦
Gemini

Web-backed Gemini integration exposed behind the local gateway boundary.

backend adapter
◌
Qwen

Qwen-compatible backend surface connected through the routing layer.

backend adapter
⌘
ChatGPT-compatible

A compatible client-facing contract can remain stable while the backend changes underneath.

compatibility layer
03 · API

Use the gateway like
a local platform.

Infera exposes familiar request shapes for developers who already work with model APIs, while keeping the service addressable on localhost.

Public surface

Representative routes exposed by the gateway.

GET /v1/modelsdiscover models
POST /v1/chat/completionsOpenAI-style chat
POST /v1/messagesAnthropic-style messages
POST /v1/responsesResponses surface
POST /v1/images/generationsimage generation surface

Request example

Point an OpenAI-compatible SDK at your local gateway and let Infera handle backend routing.

from openai import OpenAI client = OpenAI( base_url="http://127.0.0.1:8081/v1", api_key="local" ) response = client.chat.completions.create( model="your-model", messages=[{"role":"user", "content":"Hello"}] ) print(response.choices[0].message.content)
04 · Architecture

Follow the request,
layer by layer.

No magic box. The system is easiest to understand as a sequence of explicit boundaries.

01 · Clientbrowser / SDK / curl
→
02 · Gatewaylocalhost control plane
→
03 · Routermodel selection
→
04 · Adapterbackend-specific logic
→
05 · Streamnormalized response
Browser / SDK │ ▼ 127.0.0.1:8081 ── Infera gateway │ ├── model discovery ├── protocol compatibility ├── routing └── streaming │ ├────────► Gemini adapter ├────────► Qwen adapter └────────► other compatible backend surfaces Everything above the adapter boundary speaks through the local gateway contract.
05 · Trust

Open source should
show its work.

The website explains the architecture, but the repository remains the source of truth for the implementation, licenses, security notes, and setup.

⌗
Inspect the source

Read the repository before installation. Infera is designed to be understandable instead of opaque.

GitHub
◐
Local by default

The gateway's default entry point is localhost. The public page is an interface; it does not become your inference server.

localhost
※
Upstream credited

Infera retains explicit attribution to Sophomoresty and the original gemini-web2api project where applicable.

upstream
Ready when you are

AI, on your terms.

Explore the source, inspect the architecture, or open the local workspace experience.

View source