# ModelGate **Repository Path**: zmh/ModelGate ## Basic Information - **Project Name**: ModelGate - **Description**: ModelGate 基于 FastAPI 的 LLM API 代理服务器,支持多提供商管理、API Key 控制、用量监控和 Web 管理面板。 功能特性 多提供商支持:智谱、DeepSeek、Ollama、Minimax 及任意 OpenAI 兼容 API 并发限流:基于信号量的提供商级并发控制 API Key 管理:支持按 Key 限制可访问的模型 流式响应:实时流式输出,支持 token - **Primary Language**: Python - **License**: Apache-2.0 - **Default Branch**: master - **Homepage**: https://leturx.cc - **GVP Project**: No ## Statistics - **Stars**: 56 - **Forks**: 26 - **Created**: 2026-03-20 - **Last Updated**: 2026-09-27 ## Categories & Tags **Categories**: ai **Tags**: None ## README # ModelGate

ModelGate logo

ModelGate is a FastAPI-based LLM gateway for multi-provider routing, API key management, request logging, and dashboard monitoring. Designed for teams and organizations to centrally manage and distribute AI model access across departments. ## Highlights - Multi-provider routing: Zhipu, DeepSeek, Ollama, Minimax, and any OpenAI-compatible API - OpenAI-compatible proxy endpoints: `/v1/chat/completions`, `/v1/embeddings`, `/v1/models` - Anthropic-compatible proxy endpoint: `/anthropic/v1/messages` with full protocol translation (streaming, tool calls, thinking, cache_control) - Provider-less model access: API keys bind to standard models directly; call by model name, alias, or `provider/model` for explicit routing - `auto` virtual model: context-aware pool (per-item min_ctx/max_ctx) with automatic provider selection, fallback on 5xx/network errors, and route preview with exclusion reasons - Per-model context hard limit: over-limit requests are rejected up front with a native `context_length_exceeded` error, so agents (OpenCode, Claude Code) auto-compact and retry - Intent classification: auto-classify requests as coding/writing/testing/design/chat based on message content, used for smart routing and log analytics - Provider key health scoring: sliding-window (5 min) health score (0–100), prioritize healthy keys in routing - Provider key priority: manual priority per key for ordered fallback, sorted by (priority DESC, health DESC) - Layered concurrency control with queuing: concurrent requests queue up to 10s for a slot instead of failing fast; circuit-breaks to 429 (with `retry-after`) only when the wait queue is saturated - Provider multi-key support with sticky routing and key-level disable/reenable - Provider key fallback: automatically tries the next API key on 401/403/429 errors - Auto-disable provider/key on usage limit errors; scheduled auto-reenable at `reset_at + 60s` to avoid quota-window edge races - API key management with per-key model access control, expiry, regenerate, and bypass_busyness option - RBAC: users, roles, menus, and fine-grained permissions with dual auth (JWT + legacy session) - Audit logging for all admin write operations - Pricing management: per provider-model pricing with filters, batch edit, copy, sync, and CSV export - Request content logging: separate `request_contents` table for messages, response, thinking, tool_calls — lazy-loaded via Content button - Streaming request lifecycle tracking: `pending` -> `success` / `error` / `timeout` - Upstream and downstream status code logging - Model tags: assign tags to models (e.g. coding, reasoning, vision) for filtering and intent-based routing - MCP proxy: proxy remote MCP servers with API key binding, admin UI, tool sync, logging, and stats - AI-powered model recommendations and timing advice for users - AI-powered usage report generation (DOCX export with stats, trends, and fun awards) - API key time-based access rules (time windows, date ranges, weekday restrictions) - Document sharing for admin and user portal - User portal: personal stats, request history, health score, recommendations, provider availability banner, OpenCode config export - OpenCode integration: one-liner setup scripts (`irm | iex` / `curl | bash`), key-scoped model config, per-model context/output limits, reasoning-effort variants - WeChat iLink Bot integration via MCP (QR login, auto-reply, message persistence) - MinIO integration for file storage - English / Chinese i18n with Babel - Desktop and mobile admin UI with dark/light/black-gold themes - Localized static assets (no CDN dependencies) - Reverse proxy support via configurable base path - Docker Compose with Nginx reverse proxy and static file serving - Daily stats aggregation and 30-day log archiving ## Screenshots ### Admin Dashboard ![Admin Dashboard](image/admin-dashboard.png) ### Admin Monitor ![Admin Monitor](image/admin-monitor.png) ### User Dashboard ![User Dashboard](image/user-dashboard.png) ### User Report ![User Report](image/user-report.png) ### Mobile Dashboard ![Mobile Dashboard](image/mobile-dashboard.png) ## Quick Start ```bash pip install -r requirements.txt python -m app.main ``` Default local addresses: - Server: `http://localhost:8765` - Admin: `http://localhost:8765/admin/home` - User portal: `http://localhost:8765/user/login` Windows helper: `start.bat` prompts for log level and restarts the service on port 8765. ## Docker ### Docker Run ```bash docker build -t your-registry:5002/modelgate:latest . docker push your-registry:5002/modelgate:latest docker run -d --name modelgate \ -p 8765:8765 \ -e DATABASE_URL="postgresql+asyncpg://modelgate:password@host:5432/modelgate" \ -e PORT=8765 \ -e ADMIN_USERS="admin:YourPassword" \ -v /opt/modelgate/logs:/app/logs \ -v /opt/modelgate/reports:/app/reports \ -v /opt/modelgate/uploads:/app/uploads/documents \ --restart unless-stopped \ your-registry:5002/modelgate:latest ``` ### Docker Compose The repository includes a `docker-compose.yml` with ModelGate + Nginx services. Nginx handles static file serving and reverse proxying with WebSocket support. ```bash docker compose up -d ``` See [DEPLOY.md](DEPLOY.md) for full deployment instructions. ## Environment Variables | Variable | Required | Description | |----------|----------|-------------| | `DATABASE_URL` | Yes | PostgreSQL connection string | | `PORT` | No | Service port, default `8765` | | `ADMIN_USERS` | Recommended | Admin accounts, format: `user:pass,user:pass` | | `ADMIN_USERNAME` | No | Fallback admin username | | `ADMIN_PASSWORD` | No | Fallback admin password | | `LOG_LEVEL` | No | `DEBUG`, `INFO`, `WARNING`, `ERROR` | | `MINIO_ENDPOINT` | No | MinIO endpoint, default `localhost:9000` | | `MINIO_ACCESS_KEY` | No | MinIO access key | | `MINIO_SECRET_KEY` | No | MinIO secret key | | `MINIO_BUCKET` | No | MinIO bucket name, default `modelgate` | | `MINIO_SECURE` | No | Use HTTPS for MinIO, default `false` | | `ICP_NUMBER` | No | ICP filing number shown on landing page | ## Database ```sql CREATE USER "modelgate" WITH PASSWORD 'your_password'; CREATE DATABASE "modelgate" OWNER "modelgate"; ``` Schema: [`db/schema.sql`](db/schema.sql) The app performs runtime compatibility migrations on startup (e.g., adding new columns to `request_logs`). ## API ### OpenAI-compatible Endpoints - `POST /v1/chat/completions` - Chat completions (streaming and non-streaming) - `POST /v1/embeddings` - Text embeddings - `GET /v1/models` - List available models ### Anthropic-compatible Endpoint - `POST /anthropic/v1/messages` - Anthropic Messages API (streaming and non-streaming) ModelGate translates Anthropic protocol requests to OpenAI format for upstream providers, and translates responses back. Supported features: - Streaming and non-streaming responses - Tool use (function calling) with parallel tool call control - Extended thinking with signature passthrough - Cache control (`cache_control` markers on system, user, assistant, and tool_result blocks) - System prompt as structured content blocks - `inbound_protocol` tracking in request logs for protocol-level analytics ### Model Naming ```text model-name # standard model or alias (auto-select provider) provider/model # explicit provider routing auto # context-aware virtual model pool ``` Examples: `glm-5.2`, `zhipu/glm-4`, `deepseek/chat`, `auto` ## Dashboards ### Admin - `/admin/home` - Overview, realtime stats, provider availability banner, provider/key usage breakdown, trends - `/admin/config` - Providers, keys, models, bindings, routing, pricing, OpenCode setup - `/admin/api-keys` - API key management, per-key model access, expiry, regenerate - `/admin/request-logs` - Dense request log table with intent badges, content viewer, token cache breakdown - `/admin/monitor` - Composition, hotspots, response-time analysis - `/admin/reports` - AI-powered usage report generation and DOCX download - `/admin/system-config` - Outbound User-Agent management, GLM health check model, system settings - `/admin/users`, `/admin/roles` - RBAC user/role/menu management - `/admin/audit` - Audit log for all admin write operations - `/admin/mcp-servers` - MCP server management - `/admin/scheduler-tasks` - Scheduled task management with per-task cron overrides - `/admin/m` - Mobile admin dashboard ### User Portal API key holders log in at `/user/login` to access: - Personal request and token statistics (day/week/month) - My requests history with filters, pagination, and per-request token/cost breakdown - 20-minute system health score (error rate, latency, load, active users) - AI-powered model recommendations with scored reasons - AI-generated timing advice based on hourly usage patterns - Active session tracking - Model catalog with context/output limits, pricing, and multimodal info - Provider availability banner with expected restore time - OpenCode one-liner setup and configuration export ## OpenCode Integration One-liner setup merges the ModelGate provider into an existing `~/.config/opencode/opencode.jsonc` (backs up first, strips JSONC comments, never touches other providers): ```powershell # Windows (PowerShell) irm '/opencode/setup.ps1?key=sk-...' | iex ``` ```bash # macOS / Linux curl -fsSL '/opencode/setup.sh?key=sk-...' | bash ``` The generated config is scoped to the API key's model permissions and includes per-model context/output limits, an optional context hard limit, and reasoning-effort variants for thinking models. `GET /v1/models` is also key-scoped. Model visibility follows key permissions — a provider being temporarily disabled (auto circuit-break) does not remove its models from client configs. ## API Key Time-Based Access Rules API keys can be restricted by time of day, date ranges, and weekdays. Rules are validated on every request: - **Time windows** — `start_time` / `end_time` (e.g., only allow 09:00–18:00) - **Date ranges** — `start_date` / `end_date` - **Weekday filters** — restrict to specific days of the week - **Allow/deny semantics** — explicit `allowed` flag per rule ## WeChat iLink Bot (MCP) ModelGate includes an MCP (Model Context Protocol) server for WeChat iLink Bot integration at `/weixin`: - QR code login flow - Message polling, sending, and auto-reply via internal LLM proxy - Message persistence to database - Per-user context threading for conversations - See [docs/guides/weixin-mcp.md](docs/guides/weixin-mcp.md) for setup instructions ## Request Logging `request_logs` stores: API key, provider, model, tokens, latency, status, upstream/downstream HTTP status codes, client IP, user agent, intent, requested_model, actual_model, provider_key_label, and error details. Streaming requests are inserted as `pending` first, then updated to `success`, `error`, `timeout`, or `cancelled`. Logs older than 30 days are automatically archived to `request_logs_history`. A `request_logs_all` view unions both tables for transparent querying. ### Request Content (Separate Storage) Request messages, response text, thinking/reasoning, and tool calls are stored in a separate `request_contents` table, keeping `request_logs` lean for fast list queries. - **Lazy loading**: click the "Content" button in the log viewer to fetch via `GET /admin/api/logs/{id}/content` - **Cascade delete**: content rows are automatically removed when the parent log is archived or deleted ## Concurrency Control Layered semaphore-based rate control. Instead of failing fast, requests wait in queue for a slot (up to 10s by default, `MODELGATE_SEMAPHORE_ACQUIRE_TIMEOUT`); when the wait queue is saturated (>= 2x the limit), the request is rejected immediately with 429 + `retry-after` to bound queue depth under load. 1. **API key global limit** — non-`bypass_busyness` API keys are capped at 2 concurrent requests total across all providers and models 2. **API key provider-model limit** — per (api_key, provider_key, model) concurrency cap, adjustable by busyness level; `bypass_busyness` API keys skip user-side busyness concurrency limits 3. **Provider key limit** — per provider key with configurable max_concurrent 4. **System-level limit** — global concurrency with `local_rate_limited` rejection when exceeded Provider keys support sticky routing (requests from the same API key route to the same provider key). ## Key Health Scoring Each provider key has a real-time health score (0–100) based on a 5-minute sliding window: | Event | Score Impact | |-------|-------------| | Key disabled (invalid / quota exceeded) | Set to 0 | | 429/529 rate limited | -15 per event (score floors at 20 so keys can recover) | | 5xx server error | -10 per event | | 4xx client error (non-429) | -5 per event | | Successful request | +5 per 10 successes | Health levels: Excellent (90–100, green) / Good (60–89, blue) / Warning (30–59, yellow) / Critical (1–29, orange) / Unavailable (0, red). Keys are sorted by health score in `pick_api_keys`, so healthier keys are used first. ## Key Priority Provider keys support manual priority (`priority` field, default 0). `pick_api_keys` sorts by `(priority DESC, health DESC)`, enabling ordered key fallback (e.g. always try Key A first, then Key B). ## Model Routing Models can be addressed in three ways: ```text # Standard model name: auto-select provider by priority + health glm-5.2 # Alias: maps to a model, same auto-selection gpt-4o # Explicit: route to a specific provider zhipu/glm-4 # Auto virtual model: context-aware pool with fallback auto ``` When a name matches multiple providers, the routing sorts candidates by `(tag_match DESC, health DESC, priority DESC)`: 1. **tag_match**: if the model's tags include the request's intent → 1, else → 0 2. **health**: provider key health score 3. **priority**: manual priority on the provider-model binding If a provider fails with 5xx or a network error, ModelGate automatically falls back to the next candidate provider; when all providers fail, a unified "no available provider" error is returned. The `auto` virtual model routes over a configurable candidate pool with per-item `min_ctx`/`max_ctx` filtering (candidates whose context window cannot fit the request are excluded; the route preview shows exclusion reasons). ## Intent Classification Requests are auto-classified by message content into one of five intents: | Intent | Description | Badge Color | |--------|-------------|-------------| | `coding` | Programming, debugging, code review | Blue | | `writing` | Documentation, translation, editing | Amber | | `testing` | Unit tests, QA, validation | Rose | | `design` | UI/UX, wireframes, design systems | Purple | | `chat` | General conversation (default) | Gray | Classification uses keyword and code-identifier matching (no LLM call), rewritten and validated against 25k production logs. The intent is stored in `request_logs.intent` and displayed as a colored badge in the log viewer. ## Model Tags Models can be assigned tags (comma-separated) for filtering and intent-based routing: - Tags like `coding`, `reasoning`, `vision`, `flash` indicate model strengths - Tags are matched against the request intent for smart alias routing - Displayed as badges in admin config and user portal ## Provider Key Fallback When a provider has multiple API keys configured, ModelGate automatically falls back to the next key if the current one fails: - **Retryable errors**: HTTP 401 (authentication), 403 (forbidden), 429 (rate limit), 529 (overloaded) - Keys are shuffled on each request for even distribution - Sticky routing takes priority — if a sticky key is available, only that key is used - Concurrency-limited keys are skipped with a warning, trying the next available key - Fallback attempts are logged with `[KEY FALLBACK]` prefix ## Provider Auto-Disable & Reenable - When a usage limit or auth error is detected (quota exceeded, invalid key, billing deactivated), the provider or provider key is automatically disabled with a reason - Disabled state is shown in the admin dashboard and user portal (availability banner with expected restore time) - Provider keys failing auth (401/403) are auto-disabled immediately - A scheduled task reenables each disabled provider/key at `reset_at + 60s` — the 60s buffer avoids racing the upstream quota window edge - Manual reset available in admin config page ## Scheduled Tasks Cron expressions are editable per task in `/admin/scheduler-tasks`. | Task | Default Schedule | Description | |------|----------|-------------| | Daily aggregation | 00:05 | Aggregate request counts into daily/hourly stats tables | | MCP stats aggregation | 00:10 | Aggregate MCP tool usage stats | | Log archival | 00:20 | Archive request logs older than 30 days | | Request content backup | 00:30 | Export yesterday's request_contents to gzip and clean the DB | | GLM health check | 05:30 | Daily upstream health probe; notifies admins on failure | | Recommendation analysis | 08:00 | Daily model recommendation analysis | | Timeout cleanup | Every 10 minutes | Mark stale pending requests (>10 min) as `timeout` | | Busyness computation | Every 10 minutes | Compute system busyness level (1-6) | | Auto-reenable | Every 30 minutes | Reenable disabled provider keys and providers due for reset | ## Project Structure ```text modelgate/ ├── app/ # Python application package │ ├── main.py # FastAPI app entrypoint │ ├── core/ # Config, database models, i18n, path helpers │ ├── routes/ # FastAPI routers │ └── services/ # Business logic and proxy runtime ├── web/ # Runtime web assets │ ├── templates/ # Jinja2 HTML (admin/, user/, public/, components/) │ ├── static/ # CSS, JS, favicon and static frontend assets │ ├── assets/ # App image/icon assets │ └── locales/ # i18n: en, zh ├── deploy/ │ └── nginx/ # nginx.conf for Docker reverse proxy ├── db/ │ ├── schema.sql │ └── migrations/ # Database maintenance scripts ├── Dockerfile └── DEPLOY.md ``` ## Development - Python 3.10+ | FastAPI | SQLAlchemy async | PostgreSQL - Lint & format: `ruff check . && ruff format .` - Type check: `mypy app --ignore-missing-imports` - i18n compile: `pybabel compile -d web/locales` - Logs: `logs/proxy.log`, `logs/admin.log`, `logs/error.log` ## Changelog See [CHANGELOG.md](CHANGELOG.md) for release history, and [docs/CHANGELOG_2026-05_2026-08.md](docs/CHANGELOG_2026-05_2026-08.md) for a detailed optimization/bugfix deep-dive (May-Aug 2026). ## Commercial Support For production-grade cluster deployment, custom integration, or tailored feature development, contact: **minhaozhang@henngtiansoft.com** ## License Apache 2.0