# ModelGate
**Repository Path**: zmh/ModelGate
## Basic Information
- **Project Name**: ModelGate
- **Description**: ModelGate
基于 FastAPI 的 LLM API 代理服务器,支持多提供商管理、API Key 控制、用量监控和 Web 管理面板。
功能特性
多提供商支持:智谱、DeepSeek、Ollama、Minimax 及任意 OpenAI 兼容 API
并发限流:基于信号量的提供商级并发控制
API Key 管理:支持按 Key 限制可访问的模型
流式响应:实时流式输出,支持 token
- **Primary Language**: Python
- **License**: Apache-2.0
- **Default Branch**: master
- **Homepage**: https://leturx.cc
- **GVP Project**: No
## Statistics
- **Stars**: 56
- **Forks**: 26
- **Created**: 2026-03-20
- **Last Updated**: 2026-09-27
## Categories & Tags
**Categories**: ai
**Tags**: None
## README
# ModelGate
ModelGate is a FastAPI-based LLM gateway for multi-provider routing, API key management, request logging, and dashboard monitoring. Designed for teams and organizations to centrally manage and distribute AI model access across departments.
## Highlights
- Multi-provider routing: Zhipu, DeepSeek, Ollama, Minimax, and any OpenAI-compatible API
- OpenAI-compatible proxy endpoints: `/v1/chat/completions`, `/v1/embeddings`, `/v1/models`
- Anthropic-compatible proxy endpoint: `/anthropic/v1/messages` with full protocol translation (streaming, tool calls, thinking, cache_control)
- Provider-less model access: API keys bind to standard models directly; call by model name, alias, or `provider/model` for explicit routing
- `auto` virtual model: context-aware pool (per-item min_ctx/max_ctx) with automatic provider selection, fallback on 5xx/network errors, and route preview with exclusion reasons
- Per-model context hard limit: over-limit requests are rejected up front with a native `context_length_exceeded` error, so agents (OpenCode, Claude Code) auto-compact and retry
- Intent classification: auto-classify requests as coding/writing/testing/design/chat based on message content, used for smart routing and log analytics
- Provider key health scoring: sliding-window (5 min) health score (0–100), prioritize healthy keys in routing
- Provider key priority: manual priority per key for ordered fallback, sorted by (priority DESC, health DESC)
- Layered concurrency control with queuing: concurrent requests queue up to 10s for a slot instead of failing fast; circuit-breaks to 429 (with `retry-after`) only when the wait queue is saturated
- Provider multi-key support with sticky routing and key-level disable/reenable
- Provider key fallback: automatically tries the next API key on 401/403/429 errors
- Auto-disable provider/key on usage limit errors; scheduled auto-reenable at `reset_at + 60s` to avoid quota-window edge races
- API key management with per-key model access control, expiry, regenerate, and bypass_busyness option
- RBAC: users, roles, menus, and fine-grained permissions with dual auth (JWT + legacy session)
- Audit logging for all admin write operations
- Pricing management: per provider-model pricing with filters, batch edit, copy, sync, and CSV export
- Request content logging: separate `request_contents` table for messages, response, thinking, tool_calls — lazy-loaded via Content button
- Streaming request lifecycle tracking: `pending` -> `success` / `error` / `timeout`
- Upstream and downstream status code logging
- Model tags: assign tags to models (e.g. coding, reasoning, vision) for filtering and intent-based routing
- MCP proxy: proxy remote MCP servers with API key binding, admin UI, tool sync, logging, and stats
- AI-powered model recommendations and timing advice for users
- AI-powered usage report generation (DOCX export with stats, trends, and fun awards)
- API key time-based access rules (time windows, date ranges, weekday restrictions)
- Document sharing for admin and user portal
- User portal: personal stats, request history, health score, recommendations, provider availability banner, OpenCode config export
- OpenCode integration: one-liner setup scripts (`irm | iex` / `curl | bash`), key-scoped model config, per-model context/output limits, reasoning-effort variants
- WeChat iLink Bot integration via MCP (QR login, auto-reply, message persistence)
- MinIO integration for file storage
- English / Chinese i18n with Babel
- Desktop and mobile admin UI with dark/light/black-gold themes
- Localized static assets (no CDN dependencies)
- Reverse proxy support via configurable base path
- Docker Compose with Nginx reverse proxy and static file serving
- Daily stats aggregation and 30-day log archiving
## Screenshots
### Admin Dashboard

### Admin Monitor

### User Dashboard

### User Report

### Mobile Dashboard

## Quick Start
```bash
pip install -r requirements.txt
python -m app.main
```
Default local addresses:
- Server: `http://localhost:8765`
- Admin: `http://localhost:8765/admin/home`
- User portal: `http://localhost:8765/user/login`
Windows helper: `start.bat` prompts for log level and restarts the service on port 8765.
## Docker
### Docker Run
```bash
docker build -t your-registry:5002/modelgate:latest .
docker push your-registry:5002/modelgate:latest
docker run -d --name modelgate \
-p 8765:8765 \
-e DATABASE_URL="postgresql+asyncpg://modelgate:password@host:5432/modelgate" \
-e PORT=8765 \
-e ADMIN_USERS="admin:YourPassword" \
-v /opt/modelgate/logs:/app/logs \
-v /opt/modelgate/reports:/app/reports \
-v /opt/modelgate/uploads:/app/uploads/documents \
--restart unless-stopped \
your-registry:5002/modelgate:latest
```
### Docker Compose
The repository includes a `docker-compose.yml` with ModelGate + Nginx services. Nginx handles static file serving and reverse proxying with WebSocket support.
```bash
docker compose up -d
```
See [DEPLOY.md](DEPLOY.md) for full deployment instructions.
## Environment Variables
| Variable | Required | Description |
|----------|----------|-------------|
| `DATABASE_URL` | Yes | PostgreSQL connection string |
| `PORT` | No | Service port, default `8765` |
| `ADMIN_USERS` | Recommended | Admin accounts, format: `user:pass,user:pass` |
| `ADMIN_USERNAME` | No | Fallback admin username |
| `ADMIN_PASSWORD` | No | Fallback admin password |
| `LOG_LEVEL` | No | `DEBUG`, `INFO`, `WARNING`, `ERROR` |
| `MINIO_ENDPOINT` | No | MinIO endpoint, default `localhost:9000` |
| `MINIO_ACCESS_KEY` | No | MinIO access key |
| `MINIO_SECRET_KEY` | No | MinIO secret key |
| `MINIO_BUCKET` | No | MinIO bucket name, default `modelgate` |
| `MINIO_SECURE` | No | Use HTTPS for MinIO, default `false` |
| `ICP_NUMBER` | No | ICP filing number shown on landing page |
## Database
```sql
CREATE USER "modelgate" WITH PASSWORD 'your_password';
CREATE DATABASE "modelgate" OWNER "modelgate";
```
Schema: [`db/schema.sql`](db/schema.sql)
The app performs runtime compatibility migrations on startup (e.g., adding new columns to `request_logs`).
## API
### OpenAI-compatible Endpoints
- `POST /v1/chat/completions` - Chat completions (streaming and non-streaming)
- `POST /v1/embeddings` - Text embeddings
- `GET /v1/models` - List available models
### Anthropic-compatible Endpoint
- `POST /anthropic/v1/messages` - Anthropic Messages API (streaming and non-streaming)
ModelGate translates Anthropic protocol requests to OpenAI format for upstream providers, and translates responses back. Supported features:
- Streaming and non-streaming responses
- Tool use (function calling) with parallel tool call control
- Extended thinking with signature passthrough
- Cache control (`cache_control` markers on system, user, assistant, and tool_result blocks)
- System prompt as structured content blocks
- `inbound_protocol` tracking in request logs for protocol-level analytics
### Model Naming
```text
model-name # standard model or alias (auto-select provider)
provider/model # explicit provider routing
auto # context-aware virtual model pool
```
Examples: `glm-5.2`, `zhipu/glm-4`, `deepseek/chat`, `auto`
## Dashboards
### Admin
- `/admin/home` - Overview, realtime stats, provider availability banner, provider/key usage breakdown, trends
- `/admin/config` - Providers, keys, models, bindings, routing, pricing, OpenCode setup
- `/admin/api-keys` - API key management, per-key model access, expiry, regenerate
- `/admin/request-logs` - Dense request log table with intent badges, content viewer, token cache breakdown
- `/admin/monitor` - Composition, hotspots, response-time analysis
- `/admin/reports` - AI-powered usage report generation and DOCX download
- `/admin/system-config` - Outbound User-Agent management, GLM health check model, system settings
- `/admin/users`, `/admin/roles` - RBAC user/role/menu management
- `/admin/audit` - Audit log for all admin write operations
- `/admin/mcp-servers` - MCP server management
- `/admin/scheduler-tasks` - Scheduled task management with per-task cron overrides
- `/admin/m` - Mobile admin dashboard
### User Portal
API key holders log in at `/user/login` to access:
- Personal request and token statistics (day/week/month)
- My requests history with filters, pagination, and per-request token/cost breakdown
- 20-minute system health score (error rate, latency, load, active users)
- AI-powered model recommendations with scored reasons
- AI-generated timing advice based on hourly usage patterns
- Active session tracking
- Model catalog with context/output limits, pricing, and multimodal info
- Provider availability banner with expected restore time
- OpenCode one-liner setup and configuration export
## OpenCode Integration
One-liner setup merges the ModelGate provider into an existing `~/.config/opencode/opencode.jsonc` (backs up first, strips JSONC comments, never touches other providers):
```powershell
# Windows (PowerShell)
irm '/opencode/setup.ps1?key=sk-...' | iex
```
```bash
# macOS / Linux
curl -fsSL '/opencode/setup.sh?key=sk-...' | bash
```
The generated config is scoped to the API key's model permissions and includes
per-model context/output limits, an optional context hard limit, and
reasoning-effort variants for thinking models. `GET /v1/models` is also
key-scoped. Model visibility follows key permissions — a provider being
temporarily disabled (auto circuit-break) does not remove its models from
client configs.
## API Key Time-Based Access Rules
API keys can be restricted by time of day, date ranges, and weekdays. Rules are validated on every request:
- **Time windows** — `start_time` / `end_time` (e.g., only allow 09:00–18:00)
- **Date ranges** — `start_date` / `end_date`
- **Weekday filters** — restrict to specific days of the week
- **Allow/deny semantics** — explicit `allowed` flag per rule
## WeChat iLink Bot (MCP)
ModelGate includes an MCP (Model Context Protocol) server for WeChat iLink Bot integration at `/weixin`:
- QR code login flow
- Message polling, sending, and auto-reply via internal LLM proxy
- Message persistence to database
- Per-user context threading for conversations
- See [docs/guides/weixin-mcp.md](docs/guides/weixin-mcp.md) for setup instructions
## Request Logging
`request_logs` stores: API key, provider, model, tokens, latency, status, upstream/downstream HTTP status codes, client IP, user agent, intent, requested_model, actual_model, provider_key_label, and error details.
Streaming requests are inserted as `pending` first, then updated to `success`, `error`, `timeout`, or `cancelled`.
Logs older than 30 days are automatically archived to `request_logs_history`. A `request_logs_all` view unions both tables for transparent querying.
### Request Content (Separate Storage)
Request messages, response text, thinking/reasoning, and tool calls are stored in a separate `request_contents` table, keeping `request_logs` lean for fast list queries.
- **Lazy loading**: click the "Content" button in the log viewer to fetch via `GET /admin/api/logs/{id}/content`
- **Cascade delete**: content rows are automatically removed when the parent log is archived or deleted
## Concurrency Control
Layered semaphore-based rate control. Instead of failing fast, requests wait in
queue for a slot (up to 10s by default, `MODELGATE_SEMAPHORE_ACQUIRE_TIMEOUT`);
when the wait queue is saturated (>= 2x the limit), the request is rejected
immediately with 429 + `retry-after` to bound queue depth under load.
1. **API key global limit** — non-`bypass_busyness` API keys are capped at 2 concurrent requests total across all providers and models
2. **API key provider-model limit** — per (api_key, provider_key, model) concurrency cap, adjustable by busyness level; `bypass_busyness` API keys skip user-side busyness concurrency limits
3. **Provider key limit** — per provider key with configurable max_concurrent
4. **System-level limit** — global concurrency with `local_rate_limited` rejection when exceeded
Provider keys support sticky routing (requests from the same API key route to the same provider key).
## Key Health Scoring
Each provider key has a real-time health score (0–100) based on a 5-minute sliding window:
| Event | Score Impact |
|-------|-------------|
| Key disabled (invalid / quota exceeded) | Set to 0 |
| 429/529 rate limited | -15 per event (score floors at 20 so keys can recover) |
| 5xx server error | -10 per event |
| 4xx client error (non-429) | -5 per event |
| Successful request | +5 per 10 successes |
Health levels: Excellent (90–100, green) / Good (60–89, blue) / Warning (30–59, yellow) / Critical (1–29, orange) / Unavailable (0, red).
Keys are sorted by health score in `pick_api_keys`, so healthier keys are used first.
## Key Priority
Provider keys support manual priority (`priority` field, default 0). `pick_api_keys` sorts by `(priority DESC, health DESC)`, enabling ordered key fallback (e.g. always try Key A first, then Key B).
## Model Routing
Models can be addressed in three ways:
```text
# Standard model name: auto-select provider by priority + health
glm-5.2
# Alias: maps to a model, same auto-selection
gpt-4o
# Explicit: route to a specific provider
zhipu/glm-4
# Auto virtual model: context-aware pool with fallback
auto
```
When a name matches multiple providers, the routing sorts candidates by `(tag_match DESC, health DESC, priority DESC)`:
1. **tag_match**: if the model's tags include the request's intent → 1, else → 0
2. **health**: provider key health score
3. **priority**: manual priority on the provider-model binding
If a provider fails with 5xx or a network error, ModelGate automatically falls
back to the next candidate provider; when all providers fail, a unified
"no available provider" error is returned.
The `auto` virtual model routes over a configurable candidate pool with
per-item `min_ctx`/`max_ctx` filtering (candidates whose context window cannot
fit the request are excluded; the route preview shows exclusion reasons).
## Intent Classification
Requests are auto-classified by message content into one of five intents:
| Intent | Description | Badge Color |
|--------|-------------|-------------|
| `coding` | Programming, debugging, code review | Blue |
| `writing` | Documentation, translation, editing | Amber |
| `testing` | Unit tests, QA, validation | Rose |
| `design` | UI/UX, wireframes, design systems | Purple |
| `chat` | General conversation (default) | Gray |
Classification uses keyword and code-identifier matching (no LLM call),
rewritten and validated against 25k production logs. The intent is stored in
`request_logs.intent` and displayed as a colored badge in the log viewer.
## Model Tags
Models can be assigned tags (comma-separated) for filtering and intent-based routing:
- Tags like `coding`, `reasoning`, `vision`, `flash` indicate model strengths
- Tags are matched against the request intent for smart alias routing
- Displayed as badges in admin config and user portal
## Provider Key Fallback
When a provider has multiple API keys configured, ModelGate automatically falls back to the next key if the current one fails:
- **Retryable errors**: HTTP 401 (authentication), 403 (forbidden), 429 (rate limit), 529 (overloaded)
- Keys are shuffled on each request for even distribution
- Sticky routing takes priority — if a sticky key is available, only that key is used
- Concurrency-limited keys are skipped with a warning, trying the next available key
- Fallback attempts are logged with `[KEY FALLBACK]` prefix
## Provider Auto-Disable & Reenable
- When a usage limit or auth error is detected (quota exceeded, invalid key, billing deactivated), the provider or provider key is automatically disabled with a reason
- Disabled state is shown in the admin dashboard and user portal (availability banner with expected restore time)
- Provider keys failing auth (401/403) are auto-disabled immediately
- A scheduled task reenables each disabled provider/key at `reset_at + 60s` — the 60s buffer avoids racing the upstream quota window edge
- Manual reset available in admin config page
## Scheduled Tasks
Cron expressions are editable per task in `/admin/scheduler-tasks`.
| Task | Default Schedule | Description |
|------|----------|-------------|
| Daily aggregation | 00:05 | Aggregate request counts into daily/hourly stats tables |
| MCP stats aggregation | 00:10 | Aggregate MCP tool usage stats |
| Log archival | 00:20 | Archive request logs older than 30 days |
| Request content backup | 00:30 | Export yesterday's request_contents to gzip and clean the DB |
| GLM health check | 05:30 | Daily upstream health probe; notifies admins on failure |
| Recommendation analysis | 08:00 | Daily model recommendation analysis |
| Timeout cleanup | Every 10 minutes | Mark stale pending requests (>10 min) as `timeout` |
| Busyness computation | Every 10 minutes | Compute system busyness level (1-6) |
| Auto-reenable | Every 30 minutes | Reenable disabled provider keys and providers due for reset |
## Project Structure
```text
modelgate/
├── app/ # Python application package
│ ├── main.py # FastAPI app entrypoint
│ ├── core/ # Config, database models, i18n, path helpers
│ ├── routes/ # FastAPI routers
│ └── services/ # Business logic and proxy runtime
├── web/ # Runtime web assets
│ ├── templates/ # Jinja2 HTML (admin/, user/, public/, components/)
│ ├── static/ # CSS, JS, favicon and static frontend assets
│ ├── assets/ # App image/icon assets
│ └── locales/ # i18n: en, zh
├── deploy/
│ └── nginx/ # nginx.conf for Docker reverse proxy
├── db/
│ ├── schema.sql
│ └── migrations/ # Database maintenance scripts
├── Dockerfile
└── DEPLOY.md
```
## Development
- Python 3.10+ | FastAPI | SQLAlchemy async | PostgreSQL
- Lint & format: `ruff check . && ruff format .`
- Type check: `mypy app --ignore-missing-imports`
- i18n compile: `pybabel compile -d web/locales`
- Logs: `logs/proxy.log`, `logs/admin.log`, `logs/error.log`
## Changelog
See [CHANGELOG.md](CHANGELOG.md) for release history, and
[docs/CHANGELOG_2026-05_2026-08.md](docs/CHANGELOG_2026-05_2026-08.md) for a
detailed optimization/bugfix deep-dive (May-Aug 2026).
## Commercial Support
For production-grade cluster deployment, custom integration, or tailored feature development, contact:
**minhaozhang@henngtiansoft.com**
## License
Apache 2.0