Elara-Cortex Technical Report TR-2026-01
Benchmarking the Elara Route Engine against Google OR-Tools on the CVRPLIB and Time-Window Instances
ELARA·CORTEX Institute for Decision Systems and Number Theory · New Jersey · Johannesburg · hello@elara-cortex.com
Technical report, 9 June 2026
Subjects. Vehicle routing · Combinatorial optimisation · Benchmarking · Formal verification
Abstract. We compare Elara Route, a proprietary vehicle-routing solver, with published record solutions on CVRPLIB X-instances and Solomon and Gehring–Homberger time-window instances. We also run Google OR-Tools with guided local search on a subset. On nine X-instances with 100 to 512 customers, Elara records a mean gap of 3.05% in five seconds or less of single-threaded CPython. It returns a lower cost on all six instances tested with both solvers; OR-Tools gaps range from 3.6% to 16.1%. Elara matches the Z3-certified global optimum on five small instances. Across five time-window instances, its mean distance gap is 3.1%; on C101 and C1_2_1 it matches the record vehicle count at a distance gap of 0.2%. Explicit route certificates allow feasibility and cost to be checked against the public instances. A separate test replans a disturbed 200-stop, 10-vehicle day in about 1.2 ms (Section 7). We propose reporting replanning time alongside the gap at a fixed budget.
Keywords. vehicle routing; CVRPLIB; time windows; solution certificates; SMT optimality; re-optimisation under disturbance.
Introduction
Routing benchmarks usually report the gap to a published record solution. Fleet operations also depend on how quickly a plan can be revised when five stops cancel, two urgent collections appear or a vehicle fails. This report measures both quantities for Elara Route: solution cost at a stated time budget, including a comparison with Google OR-Tools, and replanning time after a disturbance.
For an instance with published record cost \(c_{\mathrm{rec}}\) and a solution of cost \(c\) found within the stated budget, the gap is the relative distance above the record,
A record solution is the lowest-cost solution any solver has published for the instance. Lower is better, and a gap of zero is the record itself.
The engine
Elara Route accepts a caller-set time budget for solving a routing problem or revising an existing plan. Its mathematical framework and implementation remain proprietary. This report evaluates the returned routes and measured execution times; Section 6 describes how readers can verify feasibility and cost without access to the solver.1
Experimental protocol
- Instances
- CVRP: nine Uchoa et al. X-instances [1] (X-n101-k25, X-n148-k46, X-n157-k13, X-n200-k36, X-n251-k28, X-n303-k21, X-n351-k40, X-n439-k37, X-n513-k21). VRPTW: Solomon 100-customer [2] instances C101, R101 and RC101, and Gehring–Homberger 200-customer [3] instances C1_2_1 and R1_2_1. Instances and record solutions come from the community-maintained PyVRP/Instances mirror.
- Distance conventions
- Official per set: integer rounded Euclidean for X (CVRPLIB EUC_2D); unrounded Euclidean with hard time windows and service times for Solomon and Gehring–Homberger.
- Baseline
- Google OR-Tools routing solver [4], guided local search, PATH_CHEAPEST_ARC start, slack vehicles, 15 to 20 seconds per instance (CVRP). Our time-window encoding returned infeasible on several instances, so it could not support a fair OR-Tools comparison. VRPTW results are compared with the published record only.
- Budgets
- Elara: five seconds or less per CVRP instance. VRPTW budgets were five seconds for C101, eight seconds each for R101 and RC101, 12 seconds for C1_2_1, and 15 seconds for R1_2_1. Runs used one thread in CPython 3.14 on a consumer Windows machine and one fixed seed. Variation across seeds remains unmeasured.
- Validation
- Every Elara solution is re-verified independently of the solver: each customer served once and no customer twice; vehicle capacity respected; time windows re-simulated arc by arc (wait when early, never late, depot closing respected); cost recomputed from the raw routes and required to match the reported value. The predicates are written out in Section 6.
Provable optimality: matching the Z3-certified global optimum
On instances small enough for an exact solver, optimality is decidable, and we certify it. We encoded the capacitated VRP as a satisfiability-modulo-theories problem and used the Z3 theorem prover [6] to compute the provable global optimum. With \(\mathcal{F}\) the set of feasible solutions of the instance and \(\tau^{*}\) the solution Elara returned, the certificate is the prover's verdict on the statement that a cheaper feasible tour exists,
so \(\tau^{*}\) is a global optimum. On every such instance the Elara engine reaches that proven optimum.
Elara against the Z3-proven global optimum on small instances, where the record solution is the optimum at this size.
| Customers | Z3 proven optimum | Elara | Verdict |
|---|---|---|---|
| 6 | 359 | 359 | optimal |
| 7 | 355 | 355 | optimal |
| 8 | 386 | 386 | optimal |
| 9 | 381 | 381 | optimal |
| 10 | 296 | 296 | optimal |
| Match | five of five | Z3-certified |
Elara matched the certified optimum on all five tested instances. In this setup, the exact method was practical only up to about 10 to 15 nodes. Larger instances are reported by their gap to the published record at the stated time budget.
Results, capacitated VRP
CVRPLIB X-instances (Uchoa et al.). The gap is distance above the published record, by equation (1). Lower is better. Elara: five seconds or less per instance, single thread CPython. OR-Tools: guided local search, 15 to 20 seconds per instance.
| Instance | Customers | Customers per route | Record | Elara gap | OR-Tools gap | Head-to-head |
|---|---|---|---|---|---|---|
| X-n101-k25 | 100 | 4.0 | 27,591 | +0.74% | +5.7% | Elara |
| X-n148-k46 | 147 | 3.2 | 43,448 | +1.64% | not run | record only |
| X-n157-k13 | 156 | 12.0 | 16,876 | +0.44% | not run | record only |
| X-n200-k36 | 199 | 5.5 | 58,578 | +2.92% | +3.6% | Elara |
| X-n251-k28 | 250 | 8.9 | 38,684 | +2.67% | not run | record only |
| X-n303-k21 | 302 | 14.4 | 21,736 | +4.32% | +8.4% | Elara |
| X-n351-k40 | 350 | 8.8 | 25,896 | +2.92% | +16.1% | Elara, 5.5 times |
| X-n439-k37 | 438 | 11.8 | 36,391 | +3.01% | +5.6% | Elara |
| X-n513-k21 | 512 | 24.4 | 24,201 | +8.79% | +15.2% | Elara, 1.7 times |
| Mean | +3.05% | at least +8% (six measured) | Elara ahead on all six measured, at three to four times less compute |
Elara returned a lower cost on all six instances tested with both solvers. The OR-Tools gap was 5.5 times Elara's gap on X-n351 and 1.7 times on X-n513 (Table 1). These comparisons apply to the configurations and budgets reported here.
Results, VRP with hard time windows
Solomon (100 customers) and Gehring–Homberger (200). Vehicles first, by the Solomon convention, then distance. The gap is on distance.
| Instance | Set | Record (vehicles · distance) | Elara (vehicles · distance) | Distance gap | Budget |
|---|---|---|---|---|---|
| C101 | Solomon C1 | 10 · 827.3 | 10 · 828.9 | +0.2% | 5 s |
| R101 | Solomon R1 | 20 · 1,637.7 | 20 · 1,657.7 | +1.2% | 8 s |
| RC101 | Solomon RC1 | 15 · 1,619.8 | 17 · 1,664.9 | +2.8% | 8 s |
| C1_2_1 | GH 200 | 20 · 2,698.6 | 20 · 2,704.6 | +0.2% | 12 s |
| R1_2_1 | GH 200 | 23 · 4,667.2 | 22 · 5,174.3 | +10.9% | 15 s |
| Serves all | 0.2% to 2.8% on the first four instances; +10.9% on R1_2_1 | under our own encoding, which is no certified head-to-head, OR-Tools returned incomplete solutions on four of five |
C101, R101 and C1_2_1 match the record vehicle counts, with distance gaps of 0.2%, 1.2% and 0.2%. RC101 uses 17 vehicles against the record's 15, with a 2.8% distance gap. R1_2_1 uses 22 against 23 and travels 10.9% farther. Because vehicle count is the first objective, the last two results should be read with both quantities in view; the summary distance range is not a vehicle-matched comparison.
Gap and route length
Across nine instances, the gap has little measured association with vehicle count (Pearson \(r = -0.06\), \(n = 9\), underpowered) and a positive association with customers per route (\(r = +0.76\), \(n = 9\); 95% confidence interval about \(0.20\) to \(0.94\)). Re-optimising the internal order of every route on X-n513, with 24 customers per route, recovered no distance. That result suggests examining the assignment of customers between routes; it does not prove that each route is globally optimal. The sample is too small to establish a general law linking route length and solver performance.
Gap by customers per route, with a budget of five seconds or less for each instance.
| Customers per route | 2 to 4 | 5 to 6 | 9 | 12 to 14 | 24 |
|---|---|---|---|---|---|
| Representative gap | 0.4% to 1.6% | 2.9% | 2.7% | 3% to 4.3% | 8.8% |
Verification without disclosure: solution certificates
Every solution in Tables 1 and 2 is published as an explicit route list (solution_certificates.json: 123 CVRP routes, 91 VRPTW routes). Re-verifying a certificate requires only the public instance file and arithmetic. Write \(\tau\) for a certificate, a set of routes \(r\), each an ordered sequence of customers beginning and ending at the depot; \(C\) for the instance's customer set; \(q_i\) for the demand of customer \(i\) and \(Q\) for the vehicle capacity; \(d_{ij}\) for the distance under the set's stated convention. A referee checks coverage and capacity,
replays the time windows arc by arc where the instance has them, with \(e_i\) and \(l_i\) the window, \(s_i\) the service time, \(t_{ij}\) the travel time and \(\pi(i)\) the stop before \(i\) on its route, so that a vehicle waits when early and is never late,
and recomputes the cost from the raw routes, which must equal the reported value:
Equations (3) to (5) are decidable from the certificate and the public instance file alone. A referee can therefore confirm every number in this report while learning nothing about how the routes were found.
Publishing route certificates allows a proprietary solver's outputs to be checked independently.
Replanning after a disturbance
The production-API test used deterministic seeds. Planning a 200-stop, 10-vehicle day took 4.6 ms. Replanning after five cancellations and three urgent insertions across five moving vehicles took about 1.2 ms. Reassigning 38 outstanding stops after a vehicle breakdown across two remaining vehicles took about 0.8 ms. These timings describe the reported run. We propose recording replanning time alongside solution quality, since both affect whether a fleet can respond to changes during the day.
Reproducibility: the public solve API
At publication, the routing API accepted CVRP and VRPTW instances at POST /v1/solve, using coordinates or a distance matrix together with demands, capacity and time windows. The response included explicit routes, cost, and verification of coverage, capacity and recomputed cost. Registered users could obtain a key through a seven-day trial at route.elara-cortex.com/developers. A comparison with another solver requires the same instance, distance convention, constraints and stated compute budget. Pickup-and-delivery, multi-depot and crew-rostering support were planned extensions at the time of this report.
Limitations
- The tables cover nine X-instances and five time-window instances. The full X and Gehring–Homberger suites, with multiple seeds, remain further work.
- The OR-Tools baseline is our configuration of a free solver; commercial solvers such as Hexaly and Gurobi are not yet in the comparison.
- The record solutions took specialised solvers hours to days [5], [8]. Nothing here claims to match them outright; the claim is the gap at a stated second-scale budget.
- All Elara timings are from single-threaded interpreted Python. Performance in a compiled deployment was not measured in this report.
Conclusion
Under the reported configurations, Elara returned lower-cost solutions than OR-Tools on all six capacitated instances tested with both solvers, using a time budget three to four times shorter. It came within 0.2% of the record distance on the clustered time-window instances while matching their vehicle counts, matched the certified optimum on five small instances, and replanned the reported disturbances in milliseconds. The results are limited to these instances, budgets and seeds. Route certificates support independent checks of feasibility and cost, using the problem conventions in the standard reference [7].
References
- Uchoa, E., Pecin, D., Pessoa, A., Poggi, M., Vidal, T., & Subramanian, A. (2017). New benchmark instances for the capacitated vehicle routing problem. European Journal of Operational Research, 257(3), 845–858.
- Solomon, M. M. (1987). Algorithms for the vehicle routing and scheduling problems with time window constraints. Operations Research, 35(2), 254–265.
- Gehring, H., & Homberger, J. (1999). A parallel hybrid evolutionary metaheuristic for the vehicle routing problem with time windows. In Proc. EUROGEN'99, 57–64.
- Perron, L., & Furnon, V. (2024). OR-Tools (v9). Google. developers.google.com/optimization
- Helsgaun, K. (2017). An extension of the Lin–Kernighan–Helsgaun TSP solver for constrained traveling salesman and vehicle routing problems. Technical Report, Roskilde University.
- de Moura, L., & Bjørner, N. (2008). Z3: An efficient SMT solver. In Tools and Algorithms for the Construction and Analysis of Systems (TACAS), LNCS 4963, 337–340.
- Toth, P., & Vigo, D. (Eds.) (2014). Vehicle Routing: Problems, Methods, and Applications (second edition). SIAM.
- Vidal, T. (2022). Hybrid genetic search for the CVRP: open-source implementation and SWAP* neighbourhood. Computers & Operations Research, 140, 105643.
Data and receipts
Instances and record solutions: CVRPLIB, Solomon and Gehring–Homberger via the community mirror. Baseline: Google OR-Tools with guided local search. Result files: results.json, lekola_validated.json, vrptw_validated.json and solution_certificates.json. The routing framework remains proprietary; the benchmark materials support verification of its outputs. Further papers are listed on the research page.
- Scope of disclosure. This report discloses results, protocol, budgets, baselines and solution certificates, which is everything needed to verify the claims. It does not disclose the method. Verification requires only the public instances and arithmetic. ↩
Cite this report
Cite as. Lekola, K. (2026). Benchmarking the Elara Route Engine against Google OR-Tools on the CVRPLIB and Time-Window Instances. Technical Report TR-2026-01, ELARA·CORTEX Institute for Decision Systems and Number Theory, New Jersey and Johannesburg. https://ai.elara-cortex.com/AI/research/tr-2026-01
@techreport{lekola2026route,
author = {Lekola, Kgomotso},
title = {Benchmarking the {Elara} Route Engine against {Google OR-Tools} on the {CVRPLIB} and Time-Window Instances},
institution = {ELARA-CORTEX Institute for Decision Systems and Number Theory},
address = {New Jersey and Johannesburg},
type = {Technical Report},
number = {TR-2026-01},
year = {2026},
month = jun,
url = {https://ai.elara-cortex.com/AI/research/tr-2026-01}
}
This version. Version 1, 9 June 2026. No DOI has been issued; cite the address above. There is no PDF: this page is the version of record, and every route list it publishes is the certificate.
Related
© 2026 ELARA·CORTEX Institute for Decision Systems and Number Theory · New Jersey · Johannesburg. All research · About Elara · Return to chat
Commentary from the Institute
About this report
A routing plan has to cope with cancelled stops, urgent collections and vehicle failures. This report measures both route quality at a fixed time budget and the time needed to replan. Published route lists let readers check feasibility and cost against the public instance files.