Hybrid Cloud + Local Architecture

Gemini Cloud Orchestration + Qwen 3.8 Local Inference

High-level planning, user interaction, and tool governance are managed by Google Gemini in the cloud, while deep code synthesis, continuous refactoring, and heavy context processing are dispatched locally to nvidia/Qwen3.8-27B-NVFP4 on your dedicated NVIDIA RTX 4500 rig.

Engine vLLM 0.31
Precision NVFP4 / FP8 KV
Context 32,768 Tok
Total Cloud Cost Avoided
$3.34 USD
≈ R$ 17.86 BRL
Total Tokens Offloaded
890,059
858,155 Prompt / 31,904 Gen
Inference Operations
71 calls
Local Tasks Executed
Prefix Cache Hit Rate
81.3%
697.8k cached / 160.4k prefilled

Cloud Quotas Remaining (Available Models)

Real-time daily capacity preserved by offloading heavy reasoning to Qwen 3.8

Daily Pool • Resets at 00:00 UTC

Google Gemini

Gemini 3.8 Flash (Orchestrator)
Orchestrator
95.3%
Remaining

OpenAI GPT

GPT-4o / o3-mini
Offloaded
97.2%
Remaining

Anthropic Claude

Claude 3.5 / 3.7 Sonnet
Offloaded
98.9%
Remaining

Cloud Cost Comparison

What you would have paid across top commercial AI providers

Token Distribution

Prefill context efficiency and generation breakdown

vLLM Telemetry

Financial Savings Matrix

Official token pricing vs your local execution cost on Twai

Provider / Model Input / 1M Output / 1M Cloud Cost (USD) Cloud Cost (BRL) Local Rig Cost Net Savings

Local Rig Infrastructure

  • Host Name: Twai (Local LAN / Tailscale)
  • Hardware Accelerator: NVIDIA RTX 4500 (24GB VRAM)
  • Tailscale Address: 100.111.95.99
  • Electricity Consumption: ~0.03 kWh (< $0.01 USD)

vLLM Engine Configuration

  • Model Architecture: nvidia/Qwen3.8-27B-NVFP4
  • Quantization Scheme: NVFP4 Weights + FP8 KV Cache
  • Context Window: 32,768 Tokens
  • Prefix Caching: Enabled (81.3% Hit Rate)