We publish a lot of numbers on this site — download speeds, latency deltas, leak test counts, kill switch success rates. Numbers are only useful if the process behind them is trustworthy, so this article is a full walkthrough of exactly how the Balumy Team builds every Security & Performance review, from lab setup to final scoring. If you’ve ever wondered why our results sometimes differ from other outlets, this is the article that explains why.
Why Methodology Transparency Matters
VPN review scores vary wildly across the internet, and a meaningful chunk of that variance has nothing to do with the products themselves — it comes down to inconsistent testing conditions, undisclosed affiliate incentives shaping conclusions, or single-run tests presented as definitive results. We built our methodology specifically to minimize those distortions, and we’re publishing it in full so readers can judge our conclusions against our process rather than just taking our word for it.
The Lab Environment
All performance testing happens on a controlled, dedicated testing rig rather than a shared office network, for one simple reason: shared networks introduce noise from other users’ traffic that can quietly bias results in either direction.
- Connection: Dedicated 1 Gbps symmetrical fiber line, isolated from any other household or office traffic during test windows.
- Hardware: A rotating set of test devices covering Windows, macOS, Linux, iOS, and Android, plus a dedicated router running OpenWRT for router-level protocol testing.
- Baseline capture: Before testing any provider, we record a fresh non-VPN baseline speed test, since ISP throughput can itself fluctuate day to day — every provider’s results are measured as a percentage of that same-day baseline, not against a fixed historical number.
Security Testing: The Five-Layer Check
Every provider goes through the same five-layer security assessment before a single speed test is run, because a fast VPN with a broken leak protection layer is not a product worth recommending regardless of its throughput numbers.
Layer 1: DNS Leak Detection
We run DNS leak checks across at least 10 distinct server locations using multiple independent detection tools cross-referenced against a custom script that logs actual outbound DNS query destinations at the network interface level. A single clean result from a browser-based tool isn’t sufficient for us to clear this layer — we want confirmation at the packet level.
Layer 2: IPv6 and WebRTC Leak Detection
Since a large share of leak incidents in the wild come from IPv6 traffic bypassing an IPv4-only tunnel, or WebRTC exposing a real IP through browser APIs, we test both explicitly, across at least three major browsers, with default browser settings (not hardened privacy configurations) to reflect how most real users actually browse.
Layer 3: Kill Switch Stress Testing
We don’t just check that a kill switch exists in the settings menu. We forcibly interrupt the VPN connection using a scripted network-drop tool, repeated 40 times per provider across different device states (active use, idle, sleep/wake transitions), and log every instance where unprotected traffic leaked through, even briefly.
Layer 4: Protocol and Cipher Verification
Using packet capture tools, we independently verify that the encryption protocol and cipher a provider advertises in its marketing is actually what’s being used on the wire, rather than taking the settings menu label at face value.
Layer 5: Audit and Transparency Review
We review publicly available independent audit reports (where they exist), infrastructure disclosures (RAM-only servers, jurisdiction, ownership transparency), and any history of past security incidents or data exposure, weighting recent, repeated audits more heavily than single one-off reports from years prior.
Performance Testing: The Four-Condition Benchmark
Once a provider clears the security layer, we move to performance testing under four distinct conditions designed to reflect real usage rather than best-case scenarios.
- Nearby server, low load: Establishes a best-case throughput ceiling.
- Long-haul server, cross-continent: Reflects realistic international streaming or access scenarios.
- Peak local hours, high load: Retests the same nearby server during evening peak usage hours in that server’s region, when real-world congestion is most likely.
- Sustained transfer test: A 30-minute continuous large file transfer to check for thermal throttling, connection drops, or gradual speed degradation that short speed-test snapshots miss entirely.
Scoring Rubric
Every provider is scored across five weighted categories that combine into a single overall Security & Performance score:
| Category | Weight | What It Measures |
|---|---|---|
| Leak Protection | 25% | DNS, IPv6, WebRTC leak test results across all tested servers |
| Kill Switch Reliability | 15% | Success rate across 40 forced-disconnect scenarios |
| Encryption Verification | 15% | Whether advertised protocol/cipher matches packet-level reality |
| Sustained Throughput | 25% | Percentage of baseline speed retained across all four load conditions |
| Audit & Transparency | 20% | Frequency, recency, and depth of independent third-party audits |
What We Deliberately Don’t Do
A few common industry practices we’ve chosen to avoid, on principle:
- Single-run testing: Every metric reported is an average of at least five separate test runs, never a single lucky (or unlucky) result.
- Vendor-supplied test environments: We never accept a provider’s own benchmarking environment or pre-configured test device. All testing happens on our own controlled hardware.
- Undisclosed timing manipulation: Some outlets test only during off-peak hours to inflate speed numbers. We explicitly include peak-hour testing in our methodology for this reason.
How Often We Retest
VPN infrastructure changes — server fleets grow, protocols get updated, audits get renewed or lapse. Every review on Balumy is scheduled for retesting on a recurring basis, and any review older than our retest window is flagged for readers so nobody mistakes a stale snapshot for a current assessment.
The Bottom Line
A methodology page isn’t the most exciting piece of content on a review site, but it’s arguably the most important one, because it’s the thing that lets you evaluate our conclusions instead of just trusting them. Every number in every Balumy Security & Performance review traces back to the process described above — controlled hardware, repeated multi-condition testing, packet-level verification, and a scoring rubric we don’t quietly adjust to fit a predetermined outcome. If you ever see a Balumy score that seems surprising, this is the page that shows exactly how we got there.
Tools We Actually Use
For readers curious about replicating pieces of this process themselves, here’s the toolkit behind the numbers:
- Packet capture: A dedicated packet analyzer running on the test network’s gateway, used to independently verify protocol and cipher claims rather than trusting app-level labels.
- DNS leak detection: Multiple independent browser-based and command-line leak testing tools, cross-referenced against each other to catch tool-specific false negatives.
- Forced disconnect scripting: A custom script that severs the underlying network interface at randomized intervals, used to stress-test kill switch engagement without relying on a provider’s own “simulate disconnect” test button, which we’ve found in some apps to behave differently from an actual real-world drop.
- Speed testing: A mix of CLI-based speed test tools and browser-based tools, run in rotation to avoid any single tool’s server selection bias skewing results in one provider’s favor.
How We Handle Conflicts of Interest
Balumy participates in standard affiliate partnerships with some of the providers we review, which is common industry practice and helps fund the cost of maintaining a dedicated testing lab. To keep that arrangement from quietly influencing conclusions, our testing and scoring process is completed in full, including the written first draft of a review, before any affiliate relationship status is factored into publication decisions. A provider that scores poorly is published as scoring poorly regardless of partnership status — and several past reviews reflect exactly that.
Handling Disagreements With Provider Claims
Occasionally our test results diverge from a provider’s own marketing claims — a speed claim that doesn’t hold up, or a leak-protection claim we can’t replicate. When that happens, we reach out to the provider directly with our specific test conditions and ask for comment or clarification before publishing, since test discrepancies are sometimes genuinely explained by configuration differences (a setting we had disabled that the provider enables by default, for instance) rather than a false claim. Where no reasonable explanation is offered, we report our measured results as-is, clearly flagged as diverging from the provider’s own claims.
Limitations We’re Upfront About
No testing lab, however careful, can fully replicate every possible real-world condition. Our results reflect a single fixed testing location’s network conditions, and readers in different geographic regions or on different ISPs may see meaningfully different absolute numbers, even if relative comparisons between providers tend to hold up reasonably well. We also can’t independently verify a provider’s internal server-side logging practices from outside their infrastructure — no reviewer can — which is why our Audit & Transparency category specifically weights third-party verification rather than claiming to have confirmed no-logs practices ourselves.
Frequently Asked Questions
How often is each review updated? Every published review is scheduled for periodic retesting, and reviews that fall outside our retest window are clearly flagged so readers know when a result may no longer reflect current infrastructure.
Do you test every plan tier a provider offers? We test the core VPN infrastructure and app, which is generally consistent across plan tiers; pricing and feature-tier differences are noted separately but don’t factor into the Security & Performance score itself.
Can a provider pay for a better score? No. Affiliate relationships fund the lab, not the scoring rubric, and the two are kept deliberately separate in our internal process.






Leave a Reply