LLM Server Cost Calculator
Estimate the electricity cost of running a local LLM server. Compare self-hosted GPU profiles against a cloud API, with live results.
Parameters
Results
Realistic Scenario
Fixed continuous-operation scenario: 60k input, 50k cached, 2k output, 90% uptime (21.6 h/day). Uses your current power, energy cost, and speed settings.
Cloud API Comparison
Compares each self-hosted profile against a cloud API for the fixed scenario (60k input / 50k cached / 2k output, 90% uptime).
Reference API: $0.25/1M input · $0.02/1M cached · $1.20/1M output
| Profile | Power | Speed (in/out) | Per Request | Monthly (self) | Monthly (API) | Savings |
|---|
Power Efficiency
| Profile | Input tok/kWh | Output tok/kWh |
|---|
How It Works
The calculator converts your GPU's power draw and token processing speeds into electricity costs using the formulas below.
Assumptions: cached tokens are processed at 1/100 the time of uncached tokens (KV-cache hit). The scenario uses 90% uptime (21.6 h/day) as a worst-case continuous-operation estimate.
Cost per hour cost_per_hour = power_w / 1000 * energy_cost Token split (cached vs uncached) cached_tokens = input_tokens * cache_rate / 100 uncached_tokens = input_tokens - cached_tokens Prompt processing time prompt_time = uncached_tokens / prompt_speed + cached_tokens / (prompt_speed * 100) Total request time request_time = prompt_time + output_tokens / decode_speed Cost per request cost_per_request = cost_per_hour * request_time / 3600 Cost per 1M input tokens cost_per_1m_input = cost_per_hour * (1_000_000 / prompt_speed) / 3600 cost_per_1m_input_cached = cost_per_hour * (1_000_000 / (prompt_speed * 100)) / 3600 Daily cost daily_cost = cost_per_request * requests_per_day Monthly cost monthly_cost = daily_cost * 30 Scenario daily cost scenario_daily = cost_per_hour * 21.6 (90% uptime)
When to Use This Calculator
This calculator helps developers, AI hobbyists, and tech writers estimate the real electricity cost of self-hosting a local LLM and compare it against cloud API pricing.
-
Estimate My Own Server's Cost
Enter your GPU's power draw, token speeds, and typical prompt/output sizes to see the real electricity cost per hour, per request, and per month.
-
Load a GPU Preset
Click a preset (RTX 4070 Ti Super, RTX 5090, RTX 5090 eco) to pre-fill realistic power and speed values, then fine-tune from there.
-
Compare Self-Hosted vs. Cloud API
See a table and chart comparing your config and the presets against a cloud API to understand monthly savings or losses of self-hosting.
-
Model a Realistic Continuous-Operation Scenario
Get a worst-case continuous-operation estimate (60k input / 50k cached / 2k output, 90% uptime) independent of your per-request inputs.
-
Understand the Formulas
Expand the 'How it works' section to see the exact cost formulas and the KV-cache assumption, so you can cite the methodology.
Frequently Asked Questions
- What does this calculator estimate?
- The electricity cost of running a local LLM server, based on your GPU's power draw, token speeds, and typical request size.
- How are cached tokens handled?
- Cached (KV-cache hit) tokens are assumed to be processed at 1/100 the time of uncached tokens, so they cost far less.
- What do the presets do?
- They pre-fill power draw and prompt/decode speeds for common GPUs so you can start from realistic numbers and adjust.
- How is the cloud API comparison calculated?
- It prices the same fixed scenario (60k input, 50k cached, 2k output) at the reference API rates and compares monthly cost against each self-hosted profile.
- Is my data sent anywhere?
- No. All calculations run in your browser and nothing is uploaded or stored.
LLM Cost Calculator Help
How to Use
- Adjust the parameters using the sliders or type exact values in the number fields
- Optionally, click a GPU preset to load realistic power and speed values
- Review the live results: cost per hour, per request, per million tokens, daily, and monthly
- Compare your config against the cloud API in the comparison table and chart
Features
- Live recalculation on every input change
- GPU presets for common cards
- Cloud API comparison with savings breakdown
- 100% client-side — no data leaves your browser
Tips
- Use the 'How It Works' section to see the exact formulas behind every number.
- The cache hit rate has a large impact on input cost — even 50% caching cuts input cost roughly in half.
- Click Reset at any time to restore all defaults and start fresh.