vcg-a40-8c-40g-16vram by Vultr
(All-cores)
(Single-core)
Specifications
Server Metadata
Vendor ID | vultr |
Name | vcg-a40-8c-40g-16vram |
Description | Cloud GPU (8 vCPUs, 40.0 GiB RAM, 740 GB NVMe, 0.3333xA40 48 GiB VRAM) |
Family | Cloud GPU |
Hw Virt | |
Average Time To Start | 69 |
Status | inactive |
Observed At | 2026-09-03T17:31:49.876020 |
Availability
| REGION / ID | SPOT | ONDEMAND |
|---|
Processor
vCPUs | 8 |
CPU Allocation | Dedicated |
CPU Cores | 4 |
CPU Speed | 2 GHz |
CPU Architecture | x86_64 |
CPU Manufacturer | Intel |
CPU L1D Cache | 32 KiB |
CPU L1D Cache Total | 256 KiB |
CPU L1I Cache | 32 KiB |
CPU L1I Cache Total | 256 KiB |
CPU L2 Cache | 4 MiB |
CPU L2 Cache Total | 16 MiB |
CPU L3 Cache | 16 MiB |
CPU L3 Cache Total | 16 MiB |
CPU Flags | fpu, vme, de, pse, tsc, msr, pae, mce, cx8, apic, sep, mtrr, pge, mca, cmov, pat, pse36, clflush, mmx, fxsr, sse, sse2, ht, syscall, nx, rdtscp, lm, constant_tsc, rep_good, nopl, xtopology, cpuid, tsc_known_freq, pni, pclmulqdq, ssse3, fma, cx16, pcid, sse4_1, sse4_2, x2apic, movbe, popcnt, tsc_deadline_timer, aes, xsave, avx, f16c, rdrand, hypervisor, lahf_lm, abm, cpuid_fault, pti, ssbd, ibrs, ibpb, fsgsbase, bmi1, avx2, smep, bmi2, erms, invpcid, xsaveopt, arat |
Ecpus | 7.9 |
Scalability | 197.5 |
System Resources and Accelerators
| MEMORY | |
|---|---|
Memory Amount | 40 GiB |
Memory Amount Actual | 40 GiB |
| GPU | |
|---|---|
GPU Count | 0.3333 |
GPU Memory Min | 16 GiB |
GPU Memory Total | 16 GiB |
GPU Manufacturer | NVIDIA |
GPU Family | Ampere |
GPU Model | A40 |
GPUs |
| STORAGE | |
|---|---|
Storage Size | 740 GB |
Storage Type | nvme ssd |
Storages |
| NETWORK | |
|---|---|
Inbound Traffic | 0 GB/month |
Outbound Traffic | 8192 GB/month |
IPv4 | 1 |
CPU and System Topology
Server Description
A cost-effective fractional GPU server featuring NVIDIA Ampere acceleration and NVMe storage for medium-scale machine learning and inference workloads.
Vultr vcg-a40-8c-40g-16vram is a Cloud GPU server designed for accelerated workloads. It features 8 dedicated Intel x86_64 vCPUs operating at 2.0 GHz, 40.0 GB of system memory, and a 740 GB NVMe SSD. Acceleration is delivered via a fractional NVIDIA Ampere A40 GPU with 16 GB of VRAM. While general CPU and memory benchmarks sit in the average tier, the instance delivers top-tier performance in LLM inference, processing 3,147.67 tokens/sec on a medium 7B model. This fractional GPU configuration provides a cost-effective entry point for machine learning and data processing tasks. It is highly suited for small to medium LLM inference, model prototyping, and database operations.
Economics
Performance
Memory Bandwidth
Compression
OpenSSL
Geekbench Single-Core
Geekbench Multi-Core
Passmark CPU Scores
| BENCHMARK | SCORE |
|---|---|
Mark | 14091 |
Compression | 174552 |
Encryption | 4817 |
Extended Instructions | 11615 |
Floating Point Maths | 31322 |
Integer Maths | 39430 |
Physics | 2083 |
Prime Numbers | 126 |
Single Threaded | 2279 |
String Sorting | 23464 |
Passmark Memory Scores
| BENCHMARK | SCORE |
|---|---|
Memory Mark | 2190 |
Database Operations | 3659 |
Memory Latency | 65 |
Memory Read Cached | 22208 |
Memory Read Uncached | 9366 |
Memory Write | 10040 |
Stress-ng Raw Scores
Stress-ng Relative Multicore Performance
LLM Inference Speed for Prompt Processing
LLM Inference Speed for Text Generation
Static Web Server
Redis
Alternatives
Servers of the Same Family
| INSTANCE | vCPUs | MEMORY | GPUs |
|---|---|---|---|
| vcg-a40-1c-5g-2vram | 1 | 5 GiB | 1⁄24 |
| vcg-a40-2c-10g-4vram | 2 | 10 GiB | 1⁄12 |
| vcg-a16-2c-8g-2vram | 2 | 8 GiB | ⅛ |
| vcg-a16-2c-16g-4vram | 2 | 16 GiB | ¼ |
| vcg-a16-3c-32g-8vram | 3 | 32 GiB | ½ |
| vcg-a40-6c-30g-12vram | 6 | 30 GiB | ¼ |
| vcg-a16-12c-128g-32vram | 12 | 128 GiB | 2 |
Similar Servers
| INSTANCE | VENDOR | vCPUs | MEMORY | GPUs |
|---|