Evidence
Engineering evidence
Fast ad decisions. Sustained execution. Measured results from the engineering behind wavebird.
Prepared by wavebird · Updated . Test dates are listed with each result.
Measured in milliseconds. Exercised over millions of requests.
Ad-path latency
28.76 ms p99
Selected result from seven local benchmark runs with a simulated exchange. March 2026.
Eight-hour memory measurement
24.3 million slot attempts
Steady-state heap trough about 2.3% above baseline in a predominantly no-fill workload. April 2026.
Sustained load
43,200 auctions
Zero failed or dropped operations in a 60-minute local load test. June 2026.
From a placement decision to a verifiable record
Request logs, decisions, browser events and settlement records describe different stages. A filled placement is an ad selected for delivery. It becomes an eligible measured impression only when the required events pass server checks. Settlement then applies the commercial terms; a test record is not a paid impression.
Product and growth
Evaluate the placement experience
Review whether ads appear in the intended moment, how no-fill behaves and whether the user can continue their AI task.
Engineering and finance
Follow the recorded outcome
Use request IDs, event status, proof records and revenue exports to distinguish delivery, eligible measurement and settlement.
March 23, 2026 · seven measured runs
28.76 ms p99 in repeated ad-path benchmarks
- Result
- End-to-end p99: 28.76 ms. Firewall p99: 0.22 ms. Simulated-exchange round-trip p99: 15.28 ms.
- What this supports
- The repeated runs demonstrated reproducible low latency across the measured decision path, from context filtering through the simulated exchange response.
- Test setup
- Seven local benchmark runs after 1,000 warmup requests, with a simulated ad exchange. Values use the selected median run for each benchmark and cover the ad path, excluding AI generation.
Measurement notes
- Samples in selected latency runs
- 12,000 firewall; 6,000 simulated-exchange; 1,200 end-to-end; 10,000 per beacon case; 2 settlement builds.
- Consistency across runs
- End-to-end p99 coefficient of variation: 9.02% across seven measured runs.
March benchmark measurements and sample counts
Measurements from March 23, 2026. Each row uses the selected median run. The settlement row reports the maximum of its two latency samples.
| Benchmark | Samples | p50 | p95 | p99 / max |
|---|---|---|---|---|
firewall_bench | 12,000 | 0.04 ms | 0.12 ms | 0.22 ms |
ssp_roundtrip_bench | 6,000 | 7.11 ms | 13.03 ms | 15.28 ms |
e2e_bench | 1,200 | 12.98 ms | 24.35 ms | 28.76 ms |
beacon_bench_seq | 10,000 | 0.37 ms | 0.93 ms | 1.39 ms |
beacon_bench_c10 | 10,000 | 5.45 ms | 8.96 ms | 10.62 ms |
beacon_bench_c50 | 10,000 | 23.83 ms | 36.46 ms | 41.76 ms |
beacon_bench_c100 | 10,000 | 52.31 ms | 73.23 ms | 78.75 ms |
settlement_bench | 2 | 832.93 ms | 887.58 ms | max 887.58 ms |
| Benchmark | Concurrency | Result |
|---|---|---|
ssp_requests_per_second | 50 | 1,364.82 requests/s |
settlement_slots_per_minute | 1 | 104,483.02 seeded slots/s |
jobs_per_second | 100 | 798.99 jobs/s |
beacons_per_second | 1 | 2,027.93 beacons/s |
The historical identifier settlement_slots_per_minute reports seeded slots per second, calculated from the batch size and elapsed seconds.
April 6–7, 2026 · eight-hour memory measurement
Memory stability across eight hours and 24.3 million slot attempts
- Result
- 24,327,994 slot attempts. The measured steady-state heap trough was about 2.3% above the post-warmup baseline. Open handles remained at 24.
- What this supports
- The memory result supported the changes to slot eviction, ledger compaction, streaming settlement and projection pruning under sustained execution.
- Test setup
- Eight hours of continuous execution against the local wrapper path at concurrency 10. This memory-stability measurement used a predominantly no-fill workload.
Measurement notes
- Workload
- 10 filled placements and 24,327,984 no-fill responses across 24,327,994 slot attempts.
- Memory measurement
- Samples every 30 seconds. Steady-state heap trough: 1.02319 times the post-warmup baseline.
April 19, 2026 · one measured run
More than 3,900 beacons per second
- Result
- Beacon throughput: 3,908.69 operations/s. Job throughput: 1,028.75 operations/s. End-to-end p99: 19.51 ms.
- What this supports
- The benchmark measured event-processing throughput alongside the decision path, providing a concrete reference for the work between an ad request and its recorded events.
- Test setup
- A local benchmark with a simulated exchange, one measured run and five warmup requests. Ad-path measurements exclude AI generation.
Measurement notes
- Additional latency results
- Firewall p99: 0.24 ms; simulated-exchange p99: 7.66 ms; sequential beacon p99: 0.72 ms; seeded settlement maximum: 0.28 ms.
- Beacon concurrency
- p99: 2.51 ms at concurrency 10, 2.52 ms at concurrency 50 and 2.11 ms at concurrency 100.
- Additional throughput
- Simulated-exchange requests: 1,486.64/s. Seeded settlement processing: 47,619.05 slots/s.
June 25, 2026 · local load test
43,200 auctions with zero failed or dropped operations
- Result
- The 60-minute soak completed 43,200 auctions at 12 auctions/s, with zero failed or dropped operations and p99 latency of 1,107.31 ms. The 24-auction/s spike and recovery also met their test criteria.
- What this supports
- The stable, spike and soak stages demonstrated sustained processing and recovery after a defined traffic spike within the tested load profile.
- Test setup
- A local load test with simulated partner services, covering a ramp, boundary probe, 30-minute stable stage, double-rate spike and 60-minute soak.
Measurement notes
- Stable stage
- 12 auctions/s for 30 minutes: 21,600 completed; no failed or dropped operations; p99 515.92 ms.
- Spike stage
- 24 auctions/s for 5 minutes: 7,200 completed; no failed or dropped operations; p99 1,639.91 ms.
- Boundary probe
- 30 auctions/s for 5 minutes: 9,000 completed; no failed or dropped operations; p99 1,145.21 ms.
Sources and methodology
Selected measurements from wavebird engineering records. Each download includes test settings, original measured values and source hashes.
Test dates identify each measurement. The March and April benchmarks include warmup and run-selection settings. The eight-hour extract records memory stability and workload counts.
Explore the integration
See how wavebird connects ad delivery, measurement and revenue in your AI product.
On this page