LLM Server Cost Calculator

Estimate the electricity cost of running a local LLM server. Compare self-hosted GPU profiles against a cloud API, with live results.

Parameters

GPU Presets
$/kWh
W
tok/s
tok/s
tok
%
tok
req

Results

Cost / Hour $0.00 Continuous operation
Per Request $0.00 0s · avg time
Monthly Cost $0.00 30 days
Cost / 1M Input (realistic) $0.00 with cache
Input tok/kWh 0 prompt efficiency
Output tok/kWh 0 decode efficiency
Cost / 1M Input (no cache) $0.00 worst case
Cost / 1M Input (all cached) $0.00 best case
Cost / 1M Output $0.00 decode only
Daily Cost $0.00 at your request rate

Realistic Scenario

Fixed continuous-operation scenario: 60k input, 50k cached, 2k output, 90% uptime (21.6 h/day). Uses your current power, energy cost, and speed settings.

Daily (90% uptime) $0.00
Monthly $0.00
Per Request $0.00 0s
Max Requests / Day 0
Input tok/kWh 0
Output tok/kWh 0

Cloud API Comparison

Compares each self-hosted profile against a cloud API for the fixed scenario (60k input / 50k cached / 2k output, 90% uptime).

Reference API: $0.25/1M input · $0.02/1M cached · $1.20/1M output

Cloud API Comparison
Profile Power Speed (in/out) Per Request Monthly (self) Monthly (API) Savings

Power Efficiency

Power Efficiency
Profile Input tok/kWh Output tok/kWh
How It Works

The calculator converts your GPU's power draw and token processing speeds into electricity costs using the formulas below.

Assumptions: cached tokens are processed at 1/100 the time of uncached tokens (KV-cache hit). The scenario uses 90% uptime (21.6 h/day) as a worst-case continuous-operation estimate.

Cost per hour
cost_per_hour = power_w / 1000 * energy_cost
Token split (cached vs uncached)
cached_tokens = input_tokens * cache_rate / 100
uncached_tokens = input_tokens - cached_tokens
Prompt processing time
prompt_time = uncached_tokens / prompt_speed + cached_tokens / (prompt_speed * 100)
Total request time
request_time = prompt_time + output_tokens / decode_speed
Cost per request
cost_per_request = cost_per_hour * request_time / 3600
Cost per 1M input tokens
cost_per_1m_input = cost_per_hour * (1_000_000 / prompt_speed) / 3600
cost_per_1m_input_cached = cost_per_hour * (1_000_000 / (prompt_speed * 100)) / 3600
Daily cost
daily_cost = cost_per_request * requests_per_day
Monthly cost
monthly_cost = daily_cost * 30
Scenario daily cost
scenario_daily = cost_per_hour * 21.6  (90% uptime)

When to Use This Calculator

This calculator helps developers, AI hobbyists, and tech writers estimate the real electricity cost of self-hosting a local LLM and compare it against cloud API pricing.

  • Estimate My Own Server's Cost

    Enter your GPU's power draw, token speeds, and typical prompt/output sizes to see the real electricity cost per hour, per request, and per month.

  • Load a GPU Preset

    Click a preset (RTX 4070 Ti Super, RTX 5090, RTX 5090 eco) to pre-fill realistic power and speed values, then fine-tune from there.

  • Compare Self-Hosted vs. Cloud API

    See a table and chart comparing your config and the presets against a cloud API to understand monthly savings or losses of self-hosting.

  • Model a Realistic Continuous-Operation Scenario

    Get a worst-case continuous-operation estimate (60k input / 50k cached / 2k output, 90% uptime) independent of your per-request inputs.

  • Understand the Formulas

    Expand the 'How it works' section to see the exact cost formulas and the KV-cache assumption, so you can cite the methodology.

Frequently Asked Questions

What does this calculator estimate?
The electricity cost of running a local LLM server, based on your GPU's power draw, token speeds, and typical request size.
How are cached tokens handled?
Cached (KV-cache hit) tokens are assumed to be processed at 1/100 the time of uncached tokens, so they cost far less.
What do the presets do?
They pre-fill power draw and prompt/decode speeds for common GPUs so you can start from realistic numbers and adjust.
How is the cloud API comparison calculated?
It prices the same fixed scenario (60k input, 50k cached, 2k output) at the reference API rates and compares monthly cost against each self-hosted profile.
Is my data sent anywhere?
No. All calculations run in your browser and nothing is uploaded or stored.