← All Assessments
Hardware Assessment • Local AI Inference

Choosing the Right GPU for
Local AI Inference

A practical comparison of NVIDIA RTX 3000, 4000, and 5000 series alongside the latest AMD RDNA 4 and Intel Arc cards. Covers pricing, specifications, and concrete recommendations for workloads ranging from simple motion detection to enterprise-scale large language model inference.

📈 Indicative Pricing 🛠 Technical Specs 🎯 6 Use Case Scenarios 🤖 CUDA / ROCm / OpenVINO 📋 Published June 2026
Executive Overview

One Metric Drives Every GPU Decision for AI: VRAM

GPU selection for local AI inference is largely a question of how much video memory (VRAM) you need, followed by how fast that memory is. Raw compute cores matter far less than they do for gaming. A card with 24GB and moderate core count will outperform a faster card with 8GB for any workload that exceeds the smaller card's memory ceiling. The hierarchy is: VRAM first, then bandwidth, then compute.

NVIDIA remains the dominant choice due to its CUDA ecosystem, which powers virtually every major AI tool from Ollama to ComfyUI to PyTorch. AMD's ROCm platform has improved substantially with RDNA 4 and now offers a viable alternative for most inference tasks. Intel Arc provides outstanding VRAM-per-pound value and is well-suited to edge inference and developer tooling via OpenVINO, though its community support lags behind NVIDIA.

Primary Decision Factor
VRAM
Video memory determines which models will even load; more VRAM equals more headroom for larger and future models
Minimum Viable Threshold
8 GB
Sufficient only for small quantized models (7B INT4) and basic CV tasks; rapidly limiting as models grow
Recommended Starting Point
12–16 GB
Comfortable for image generation, 7B–14B LLMs, and day-to-day AI developer workflows
CUDA Ecosystem Share
Dominant
Nearly all consumer AI tools, models, and community guides assume NVIDIA CUDA; other platforms require more setup
What This Assessment Covers Eighteen consumer and prosumer GPUs across five product families are compared on VRAM, memory bandwidth, power draw, AI software ecosystem maturity, and indicative UK retail pricing. Six real-world scenarios are mapped to specific hardware recommendations. Professional data centre cards (NVIDIA A100, H100, RTX 6000 Ada) are out of scope, though noted where relevant.
Use Case Guide

Match Your Workload to the Right Hardware

Identify the scenario that best matches your primary use, then use the budget, balanced, and performance picks as your shortlist. Prices are indicative UK retail; check current listings before purchasing.

📷

Motion and Face Detection

Security cameras, access control, edge CV

💾 VRAM needed: 2–4 GB
Budget Intel Arc B580 12GB ~£240
Balanced RTX 3060 12GB ~£220
Performance RTX 4060 8GB ~£250
This workload (YOLOv8, MediaPipe, OpenCV DNN, MobileNet) is extremely light. Any modern GPU handles it easily. The recommended cards are chosen for software compatibility and efficiency rather than raw power. Intel Arc is particularly strong here via OpenVINO, which is purpose-built for edge inference. A dedicated older GTX 1060 or GTX 1650 also handles this if cost is the primary concern.
💻

Web and App Development

AI coding assistants, occasional local model testing

💾 VRAM needed: 8 GB (12–16 GB recommended)
Budget RTX 4060 8GB ~£250
Balanced Arc A770 16GB or RTX 4060 Ti 16GB ~£270 / ~£400
Performance RTX 4070 Super 12GB ~£510
Cloud AI coding tools (GitHub Copilot, Cursor, Windsurf) run inference in the cloud and do not require local GPU power. A GPU here primarily supports local model testing, CUDA-accelerated build tools, and GPU-based browser profiling. The Arc A770 offers 16GB GDDR6 at very low cost, making it exceptional value for the developer who wants local model headroom without the price of a mid-range NVIDIA card.
🤖

Vibe Coding

Running 7B–34B coding LLMs locally as primary AI assistant

💾 VRAM needed: 8 GB minimum, 16 GB recommended
Budget RTX 4060 Ti 16GB ~£400
Balanced RTX 4070 Super 12GB ~£510
Performance RTX 4070 Ti Super 16GB ~£720
Models such as Qwen2.5-Coder 14B, DeepSeek-Coder-V2 Lite, and Codestral 7B fit within 12–16GB at Q4 quantization. Token generation speed (tokens per second) depends heavily on memory bandwidth: the RTX 4060 Ti 16GB has lots of VRAM but a narrow 128-bit bus (288 GB/s), which limits speed noticeably. The RTX 4070 Super has a faster 192-bit bus (504 GB/s) and generates tokens approximately 75% faster despite having less VRAM. For vibe coding where latency matters, bandwidth wins over raw capacity.
🎨

Image Generation

ComfyUI, Automatic1111, SDXL, Flux.1 Schnell

💾 VRAM needed: 8 GB (SD 1.5), 12 GB (SDXL), 16–24 GB (Flux)
Budget RTX 4060 Ti 16GB ~£400
Balanced RTX 4070 Ti Super 16GB ~£720
Performance RTX 4090 24GB ~£1,650
Stable Diffusion 1.5 runs on 8GB. SDXL is comfortable at 12GB and workable at 8GB with tiling. Flux.1 Dev in full FP16 precision requires 24GB; in NF4 quantization it is usable at 16GB but with a speed penalty. For regular SDXL and LoRA workflows, the RTX 4060 Ti 16GB delivers excellent value. If Flux.1 or video generation is the goal, 16GB becomes a ceiling: budget for 24GB.
🎞

Professional Image and Video

Flux.1 Dev native, video generation, multi-model pipelines

💾 VRAM needed: 24 GB minimum, 32 GB ideal
Budget RX 7900 XTX 24GB (ROCm) ~£760
Balanced RTX 4090 24GB ~£1,650
Performance RTX 5090 32GB ~£2,100
Flux.1 Dev at native FP16 precision, video generation models (CogVideoX 5B, Wan 2.1), and multi-model pipelines all demand 24GB or more. The RX 7900 XTX offers 24GB GDDR6 at a substantially lower price than the RTX 4090 and ROCm support in PyTorch is now mature enough for ComfyUI and diffusion pipelines. For production workflows where CUDA compatibility is critical, the RTX 4090 remains the standard. The RTX 5090 32GB is the first consumer card to exceed the 24GB barrier and is the right choice for video generation becoming a primary business activity.
📊

Big Data and LLM Inference

30B–70B parameter models, RAG, embedding at scale

💾 VRAM needed: 24 GB (30B INT4), 40–48 GB (70B INT4)
Consumer RTX 5090 32GB ~£2,100
Dual GPU RTX 4090 x2 (NVLink) ~£3,300
Cloud RunPod / Lambda Labs A100 per hour
Llama 3.1 70B at Q4 quantization requires approximately 40GB VRAM. A single RTX 5090 32GB handles Llama 3.1 34B and most 30B-class models in INT4; for the full 70B tier, two RTX 4090s via NVLink (48GB combined) or a cloud A100/H100 instance is required. For intermittent workloads, cloud GPU rental is almost always more cost-effective than building dedicated local hardware above the RTX 4090 tier.
Top Picks

Seven Cards Worth Shortlisting

Across all scenarios and budgets, these are the cards that represent the most defensible choices. Every other card in the spec table either costs more for similar results, or serves a narrower use case.

Best VRAM Value
Intel Arc A770 16GB
~£270 indicative
16 GB GDDR6 • 560 GB/s • 225W
The cheapest path to 16GB VRAM by a significant margin. Intel OpenVINO and IPEX provide solid inference support. Best for developers who want local model headroom without the cost of an Ada or Blackwell card. CUDA tools require extra setup via workarounds.
Intel Arc
Best Used-Market Pick
RTX 3060 12GB
~£220 indicative
12 GB GDDR6 • 360 GB/s • 170W
Full CUDA support, mature driver stack, 12GB VRAM at one of the lowest price points in the CUDA ecosystem. 3rd-gen Tensor Cores are slower than Ada, but entirely functional for inference workloads. The right choice when budget is the primary constraint and CUDA compatibility is required.
Best Budget CUDA
Best Value 16 GB CUDA
RTX 4060 Ti 16GB
~£400 indicative
16 GB GDDR6 • 288 GB/s • 165W
The cheapest Ada Lovelace card with 16GB VRAM. The narrow 128-bit memory bus (288 GB/s) limits LLM token generation speed noticeably; this card is better suited to image generation and model loading than for fast LLM inference. VRAM-constrained buyers on a mid-range budget.
Max VRAM, Min Cost
Best Mid-Range Balanced
RTX 4070 Super 12GB
~£510 indicative
12 GB GDDR6X • 504 GB/s • 220W
The sweet spot of the Ada generation for AI use. Fast 192-bit bus delivers nearly double the inference throughput of the RTX 4060 Ti 16GB despite having less VRAM. Handles SDXL, 7B coding models, and most developer workflows with excellent responsiveness. Slightly limited for Flux.1 or 13B+ models.
Best All-Round Mid
Best Premium Ada Pick
RTX 4070 Ti Super 16GB
~£720 indicative
16 GB GDDR6X • 672 GB/s • 285W
Combines 16GB VRAM with a wide 256-bit bus and Ada Lovelace Tensor Cores. Runs 13B LLMs quickly, generates SDXL images fast, handles Flux.1 in NF4 quantization, and manages basic video generation. The ceiling of sensible mainstream AI hardware for developers and serious enthusiasts.
Speed and Capacity
Best Performance GPU
RTX 4090 24GB
~£1,650 indicative
24 GB GDDR6X • 1,008 GB/s • 450W
The reference-class GPU for local AI research and professional creative work. 24GB VRAM, the fastest consumer bandwidth in the Ada generation, and a mature 4th-gen Tensor Core stack. Handles Flux.1 Dev natively, 30B LLMs in INT4, and all current diffusion pipelines. The de facto standard for serious AI practitioners.
Reference Standard
Maximum Future-Proofing
RTX 5090 32GB
~£2,100 indicative
32 GB GDDR7 • 1,792 GB/s • 575W
The only consumer card with 32GB GDDR7. NVIDIA Blackwell introduces 5th-gen Tensor Cores with FP4 and FP6 support, enabling new quantization formats that meaningfully expand what fits in memory. Handles 70B INT4 models with offloading, all current video generation models, and positions the buyer ahead of the next 18–24 months of model size growth.
Maximum Headroom
Note on RTX 5000 Series Availability The RTX 5070, 5070 Ti, 5080, and 5090 launched in early 2025. Street prices have in many markets exceeded MSRP due to constrained supply. Verify current availability and pricing from UK retailers before committing. The RTX 4000 series remains widely available and often represents better value per pound where street prices for 5000-series cards remain elevated.
Technical Analysis

Specifications, Ecosystem, and Pricing

The sections below cover the technical detail behind the recommendations: how to read GPU specs for AI inference, the full comparison table, AI framework ecosystem maturity, and a visual pricing analysis.

VRAM Tiers

Understanding VRAM Requirements

GPU memory (VRAM) is the hard ceiling for AI inference. When a model's weights exceed available VRAM, the workload must spill to system RAM, reducing performance dramatically, or simply fail to load. Knowing the VRAM tier you need is the single most important step in GPU selection.

8 GB
SD 1.5, 7B Q4 quantized, basic detection
12 GB
SDXL, 7B FP16, 13B Q4, Flux Schnell (with offload)
16 GB
Flux.1 Dev NF4, 13B FP16, 34B Q4, SDXL multi-LoRA
24 GB
Flux.1 Dev FP16, video generation, 30B Q4, multi-model pipelines
32 GB
70B Q4 (with partial offload), large embedding batches, long-context RAG

🔈 Memory Bandwidth: The Speed Multiplier

Once the model fits in VRAM, inference speed is primarily determined by memory bandwidth: how fast the GPU can move weights from VRAM to compute cores. This is measured in gigabytes per second (GB/s). For LLM inference specifically, token generation is almost entirely memory-bandwidth bound; adding CUDA cores does not help if the bus is the bottleneck.

This explains the performance gap between the RTX 4060 Ti 16GB (288 GB/s) and the RTX 4070 Super 12GB (504 GB/s). Despite having more VRAM, the 4060 Ti generates tokens roughly 40–50% slower in typical LLM workloads. For image generation, bandwidth matters less because diffusion steps are more compute-intensive; here the VRAM advantage of the 4060 Ti becomes relevant.

Key bandwidth thresholds for LLM inference

  • Under 300 GB/s: noticeably slow token generation for 7B+ models; viable for edge inference and low-volume use
  • 300–600 GB/s: comfortable for 7B models, usable for 13B; the mainstream range
  • 600–1,000 GB/s: fast inference on 7B–30B models; this tier includes the RTX 4070 Ti Super, 4080 Super, and 4090
  • Over 1,000 GB/s: only the RTX 4090 (1,008 GB/s) and RTX 5090 (1,792 GB/s) in the consumer market; near-professional throughput

⚡ Power Consumption and Thermal Considerations

AI inference is a sustained workload, not a burst load. Unlike gaming, where GPU usage fluctuates, running an LLM or diffusion model sustains near-maximum power draw for the duration of the inference task. This means TDP (thermal design power) directly determines electricity cost and cooling requirements.

A RTX 4090 running at 450W sustained for 8 hours per day consumes approximately 3.6 kWh daily. At UK electricity rates of around 24p per kWh, this is approximately 86p per day or around £315 per year, before any other system components. The RTX 5090 at 575W raises this further. Cards in the 165–250W range (RTX 4060 Ti, RTX 4070 Super, RTX 5070) are substantially cheaper to run.

Ensure your PC case provides adequate airflow for sustained inference loads. Many coolers rated for gaming loads underperform when GPU utilisation remains at 100% for extended periods.

Technical Specifications

Full GPU Specification Comparison

Specifications are as published by the respective manufacturers. UK pricing is indicative and based on observed street prices through mid-2025; current retail prices may differ. Always verify before purchasing. Cards highlighted with a warning on VRAM carry real-world limitations for typical AI workloads.

How to Read This Table For AI inference, prioritise: (1) VRAM, (2) Memory Bandwidth, (3) AI Framework support. CUDA core counts and clockspeeds are less important than they are for rendering or gaming.

NVIDIA RTX 3000 Series (Ampere, 2020–2021)

GPU VRAM Bandwidth CUDA Cores Tensor Gen TDP ~UK Price AI Suitability
RTX 3060 12GB 12 GB GDDR6 360 GB/s 3,584 3rd Gen 170W ~£220 Good value CUDA entry
RTX 3070 8GB 8 GB GDDR6 448 GB/s 5,888 3rd Gen 220W ~£230 VRAM limits modern use
RTX 3080 10GB 10 GB GDDR6X 760 GB/s 8,704 3rd Gen 320W ~£330 Fast but tight VRAM
RTX 3090 24GB 24 GB GDDR6X 936 GB/s 10,496 3rd Gen 350W ~£580 Excellent used-market value

NVIDIA RTX 4000 Series (Ada Lovelace, 2022–2024)

GPU VRAM Bandwidth CUDA Cores Tensor Gen TDP ~UK Price AI Suitability
RTX 4060 8GB 8 GB GDDR6 272 GB/s 3,072 4th Gen 115W ~£250 Constrained; dev tooling only
RTX 4060 Ti 16GB 16 GB GDDR6 288 GB/s 4,352 4th Gen 165W ~£400 Good VRAM; slow inference bus
RTX 4070 Super 12GB 12 GB GDDR6X 504 GB/s 7,168 4th Gen 220W ~£510 Best mid-range balanced
RTX 4070 Ti Super 16GB 16 GB GDDR6X 672 GB/s 8,448 4th Gen 285W ~£720 Excellent; top mainstream pick
RTX 4080 Super 16GB 16 GB GDDR6X 736 GB/s 10,240 4th Gen 320W ~£930 Strong; limited VRAM edge vs 4070 TiS
RTX 4090 24GB 24 GB GDDR6X 1,008 GB/s 16,384 4th Gen 450W ~£1,650 Reference-class AI GPU

NVIDIA RTX 5000 Series (Blackwell, 2025)

GPU VRAM Bandwidth CUDA Cores Tensor Gen TDP ~UK Price AI Suitability
RTX 5070 12GB 12 GB GDDR7 672 GB/s 6,144 5th Gen 250W ~£600 Faster than 4070S; same VRAM ceiling
RTX 5070 Ti 16GB 16 GB GDDR7 896 GB/s 8,960 5th Gen 300W ~£800 Strong update on 4070 TiS
RTX 5080 16GB 16 GB GDDR7 960 GB/s 10,752 5th Gen 360W ~£1,050 Fast but VRAM capped at 16GB
RTX 5090 32GB 32 GB GDDR7 1,792 GB/s 21,760 5th Gen 575W ~£2,100 Maximum consumer AI capability

AMD Radeon (RDNA 3 and RDNA 4)

GPU VRAM Bandwidth Stream Processors Architecture TDP ~UK Price AI Suitability
RX 7900 XTX 24GB 24 GB GDDR6 960 GB/s 6,144 (96 CU) RDNA 3 355W ~£760 Good ROCm support; CUDA gap remains
RX 9070 XT 16GB 16 GB GDDR6 640 GB/s 4,096 (64 CU) RDNA 4 304W ~£550 Best AMD value; ROCm mature on RDNA 4

Intel Arc (Alchemist and Battlemage)

GPU VRAM Bandwidth Compute Units Architecture TDP ~UK Price AI Suitability
Arc A770 16GB 16 GB GDDR6 560 GB/s 32 Xe-cores (4,096 ALU) Alchemist 225W ~£270 Excellent VRAM/price; OpenVINO inference
Arc B580 12GB 12 GB GDDR6 456 GB/s 20 Xe2-cores (2,560 ALU) Battlemage 190W ~£240 Best VRAM-per-pound; edge inference ideal
RTX 3070, 3080, and 4060: Cards Intentionally Omitted from Top Picks The RTX 3070 8GB and RTX 4060 8GB are passed over in most AI scenarios due to the 8GB VRAM ceiling, which is increasingly limiting. The RTX 3080 10GB has better bandwidth but also sits below the 12GB threshold. These cards remain capable gaming GPUs; for AI inference, directing budget toward a card with 12GB or more is almost always the better decision.
Software Ecosystem

CUDA, ROCm, and Intel OpenVINO: Platform Comparison

Your GPU is only as useful as the software that runs on it. The AI framework ecosystem is the most underweighted factor in GPU purchasing decisions, and it materially affects how much friction you will encounter setting up and maintaining AI workloads.

🟢 NVIDIA CUDA
  • PyTorch, TensorFlow, and JAX all default to CUDA
  • Ollama, llama.cpp, and llama-server have native CUDA support
  • ComfyUI, Automatic1111, and InvokeAI are built and tested on CUDA first
  • Community guides, tutorials, and Stack Overflow answers almost exclusively assume CUDA
  • NVIDIA cuDNN, TensorRT, and cuBLAS add significant inference optimisation
  • Tensor Core generations (3rd through 5th) each add new precision formats; FP4 and FP6 in Blackwell are material for quantised LLMs
Ecosystem Rating: Dominant Standard
🔴 AMD ROCm
  • PyTorch ROCm officially supports RDNA 3 (RX 7000) and RDNA 4 (RX 9000) series
  • ComfyUI and llama.cpp work on ROCm with modest configuration overhead
  • Many community extensions and custom nodes test only on CUDA; ROCm may require patching
  • TensorRT and cuDNN have no AMD equivalents; ROCm performance optimisation is less mature
  • Some models with CUDA-specific kernels (FlashAttention, Triton) need fallback paths on ROCm
  • AMD MIOpen provides convolution and attention acceleration for RDNA 4
Ecosystem Rating: Capable, Some Friction
🟠 Intel OpenVINO and IPEX
  • OpenVINO is purpose-built for inference; excellent INT8 and BF16 performance on Arc
  • Intel Extension for PyTorch (IPEX) adds Arc GPU acceleration to PyTorch workloads
  • llama.cpp SYCL backend supports Intel Arc; Ollama provides limited Arc support
  • ComfyUI and Automatic1111 require additional configuration and may lack feature parity
  • Training workloads are not well-supported; Intel Arc is an inference platform
  • Community documentation and issue resolution assumes CUDA; Arc problems take longer to diagnose
Ecosystem Rating: Inference Only; Limited Community

🔌 Practical Ecosystem Impact

For a developer getting started with local AI for the first time, the CUDA ecosystem advantage is significant. A new NVIDIA GPU will work with Ollama, ComfyUI, and most HuggingFace examples without any additional configuration. The same workloads on AMD or Intel require reading platform-specific documentation, adjusting environment variables, and occasionally debugging library mismatches.

This gap has narrowed considerably through 2024 and 2025. AMD RDNA 4 with ROCm 6.x is a genuinely usable AI platform for most diffusion and LLM inference workloads. Intel Arc is the stronger choice for production inference endpoints and edge deployments via OpenVINO, but less suited to the exploratory and rapidly-evolving community tooling ecosystem.

The ROCm and OpenVINO ecosystems are strategically important. Both AMD and Intel are investing heavily to close the gap with CUDA, and both platforms will improve over the 2026–2028 cycle. Buyers comfortable with some initial configuration overhead can save significantly by choosing AMD, particularly the RX 9070 XT 16GB at approximately £550 versus the RTX 4070 Ti Super at £720.

Pricing Analysis

Value for VRAM: Where Each Card Sits

The charts below show approximate indicative UK street pricing and the derived cost per gigabyte of VRAM. Cost per GB of VRAM is a useful shorthand for AI purchasing decisions: a low number means you are getting more memory headroom per pound spent.

VRAM by GPU (Gigabytes)

Colour indicates GPU series. 8 GB cards are highlighted in amber as they are below the recommended floor for most AI workloads. Indicative data; verify before purchasing.

RTX 3000 (Ampere)
RTX 4000 (Ada)
RTX 5000 (Blackwell)
AMD
Intel Arc

Memory Bandwidth (GB/s)

Memory bandwidth drives LLM token generation speed. Cards below 300 GB/s will be noticeably slow for real-time inference on 7B+ parameter models.

Indicative UK Street Price (GBP)

Approximate street prices based on observed UK retail through mid-2025. Prices vary by retailer, availability, and over time. RTX 5000-series cards may exceed MSRP where supply is constrained. Always confirm current pricing.

RTX 3000 (Ampere)
RTX 4000 (Ada)
RTX 5000 (Blackwell)
AMD
Intel Arc
The RTX 4080 Super Problem The RTX 4080 Super 16GB sits at approximately £930, offering 736 GB/s bandwidth and 16GB VRAM. The RTX 4070 Ti Super delivers 672 GB/s and the same 16GB VRAM for £720 or less. The approximately 9% bandwidth gain of the 4080 Super does not justify the 29% price premium for AI inference tasks. Only buyers who also need the higher rasterisation performance for gaming or 3D rendering will find the 4080 Super the right choice.
Final Recommendation

The Right Card for Each Use Case

Summarising all technical and pricing analysis into a practical decision table. Each recommendation is based on the optimal balance of VRAM, bandwidth, ecosystem maturity, and price for the stated use case.

📷

Motion Detection and Face Recognition

Edge CV, security systems, basic object detection. The workload is trivial for any modern GPU. Choose based on power efficiency and CUDA availability.

● Intel Arc B580 12GB
~£240 • 190W
◉ RTX 3060 12GB
~£220 (CUDA required)
💻

Web and App Development with AI Tools

IDE AI assistants, local model testing, AI-assisted debugging. 8GB acceptable but 16GB is more comfortable for pulling and running model variants.

● Intel Arc A770 16GB
~£270 • 16GB at lowest cost
◉ RTX 4060 Ti 16GB
~£400 (full CUDA)
🤖

Vibe Coding with Local LLMs

Running 7B–14B coding models (Qwen2.5-Coder, DeepSeek-Coder) as primary AI development assistant. Bandwidth matters more than VRAM capacity here.

● RTX 4070 Super 12GB
~£510 • 504 GB/s
◉ RTX 4070 Ti Super 16GB
~£720 (16GB + fast bus)
🎨

Image Generation (Enthusiast)

SDXL, ComfyUI workflows, Flux.1 Schnell, multiple LoRA combinations. 16GB covers most workflows; 24GB needed for Flux.1 Dev at full precision.

● RTX 4070 Ti Super 16GB
~£720 • balanced for SDXL+
◉ RTX 4090 24GB
~£1,650 (Flux.1 Dev native)
🎞

Professional Image and Video Production

Flux.1 Dev at full precision, CogVideoX, Wan 2.1 video generation, commercial multi-model pipelines. 24GB is a practical minimum; 32GB removes friction.

● RTX 4090 24GB
~£1,650 • established standard
◉ RTX 5090 32GB
~£2,100 (future-proofed)
📊

Big Data Analytics and Enterprise LLM Inference

Running 30B–70B parameter models, large-scale RAG, embedding pipelines. For 70B+ models, cloud GPU instances are usually more cost-effective than local hardware.

● RTX 5090 32GB
~£2,100 • 30B–34B native
◉ Cloud A100 / H100
Per-hour (70B+ workloads)

👉 The Short Version

If budget is the primary constraint, buy the card with the most VRAM you can afford that uses NVIDIA CUDA. For most users that means the RTX 3060 12GB (used), the RTX 4060 Ti 16GB, or the RTX 4070 Super 12GB depending on whether you prioritise VRAM or bandwidth. If you can stretch to the RTX 4070 Ti Super 16GB, that is the single best all-round AI GPU for developers and creative professionals who do not need 24GB. The RTX 4090 is the ceiling for serious local AI; the RTX 5090 is for those who need to push beyond it.

3060 12GB: Best Budget CUDA 4060 Ti 16GB: Most VRAM per £ 4070 Super: Best Balanced 4070 Ti Super: Best All-Rounder 4090: Reference Class 5090: Maximum Arc A770: Best Intel Value RX 9070 XT: Best AMD Value
Pricing Disclaimer All UK prices shown in this assessment are indicative order-of-magnitude figures based on observed retail prices through mid-2025. GPU prices fluctuate with supply, demand, currency movements, and product lifecycle. Figures should be treated as a relative positioning guide only. Verify current pricing from UK retailers (Scan, Overclockers, AWD-IT, Amazon UK) before purchasing. This assessment does not constitute a commercial offer or guaranteed pricing.