pepedocs

Benchmarks

pepe is built to cost less than the server it tests. It is measured against oha, vegeta, wrk and k6 on every release, with the suite in bench/, against a local server that answers from memory, so the client is the cost being measured. CPU milliseconds per 1,000 requests, peak memory, and the requests per second each tool reported.

Apple M4 Pro

Single-thread numbers for pepe, where the loopback tops out near 175k requests a second:

Workloadpepeohavegeta
GET, 64 connections6.2 ms · 9 MB · 160k req/s23.0 ms · 78 MB · 161k req/s90.4 ms · 26 MB · 76k req/s
GET, 1,000 connections6.0 ms · 12 MB · 163k req/s21.6 ms · 101 MB · 147k req/s52.7 ms · 92 MB · 131k req/s
HTTPS, 64 connections6.8 ms · 12 MB · 142k req/s24.8 ms · 60 MB · 156k req/s69.8 ms · 30 MB · 90k req/s

Linux

A 4-vCPU arm64 VM, the musl build that is released, with wrk on one thread beside it:

Workloadpepewrkoha
GET, 64 connections2.4 ms · 4.0 MB · 398k req/s3.5 ms · 4.5 MB · 282k req/s6.8 ms · 67 MB · 305k req/s
GET, 1,000 connections3.2 ms · 5.8 MB · 296k req/s4.4 ms · 7.7 MB · 228k req/s5.6 ms · 73 MB · 183k req/s
HTTPS, 64 connections3.3 ms · 6.0 MB · 284k req/s4.5 ms · 10.8 MB · 218k req/s8.1 ms · 42 MB · 250k req/s
10 million requests, 256 connections2.9 ms · 4.5 MB · 343k req/s3.3 ms · 4.6 MB · 304k req/s6.5 ms · 2,404 MB · 316k req/s

How

A share-nothing engine: each sending thread keeps its own connections and reads and writes them itself. The request is bytes made before the run; the response is parsed where it was read; nothing is allocated for a request, and a thread's connections share one read buffer. Proxies, redirects and flows go through a general client instead, and the benchmark's README says where pepe is level rather than ahead (a slow target at 1,000 connections, 16 KB bodies over TLS), what a glibc build changes, and how k6 does.

Where a target can take more than one thread sends, pepe says so, in the footer, the verdict and generator.peak_busy_percent, and --threads auto adds them (Load testing: threads).

Every workload, the profiles and the method are in bench/README.md; a benchmark gate runs on every release pull request.

Edit this page on GitHub