Benchmarks
Farce also aims to be fast and efficient, aiming for anywhere between a minimal overhead to outperforming other options.
Any numbers quoted here are to be taken with a grain of salt:
- They are based on micro-benchmarks, which may not reflect real-world performance.
- They are momentary snapshots. Gems and Ruby implementations are constantly evolving, so these numbers may not be accurate in the future.
- Concrete numbers are largely measured on a local machine, which may not reflect your deployment environment.
You should measure the performance of your own application under realistic conditions.
Map performance
A shared map is a very common data structure for tracking state. While the read and write performance is hopefully not the bottleneck for your application.
| Map implementation | Read | Write | Notes |
|---|---|---|---|
Hash |
fastest | fastest | Not ractor-shareable (when mutable), not thread-safe |
Concurrent::Hash |
1.1x slower | 1.1x slower | Not ractor-shareable |
Farce::Strict::Map |
1.2x slower | 1.1x slower | |
Farce::Unshared::Map |
1.2x slower | 1.1x slower | Not ractor-shareable |
Farce::Map |
1.3x slower | 1.1x slower | |
Concurrent::Map |
1.4x slower | 3.2x slower | Not ractor-shareable |
Hash + Mutex |
2.9x slower | 2.9x slower | Not ractor-shareable |
Ractor::LockHash |
2.9x slower | 3.4x slower | Not fiber-friendly |
Ratomic::Map |
3.1x slower | 2.7x slower | Breaks isolation, not fiber-friendly |
Ractor::KeyLockHash |
3.2x slower | 2.5x slower | Not fiber-friendly |
Farce::LRUMap |
3.4x slower | 7.6x slower | Automatic eviction |
Farce::LFUMap |
3.5x slower | 7.6x slower | Automatic eviction |
Farce::Unsafe::TreeMap |
3.7x slower | 4.1x slower | Ordered entries, not thread-safe, not ractor-shareable |
HashWithIndifferentAccess |
4.5x slower | 5.3x slower | Not ractor-shareable |
RactorSafe::HashMap |
4.8x slower | 4.8x slower | |
Farce::TreeMap |
5.5x slower | 40x slower | Ordered entries |
Farce::LeaseMap |
30x slower | 40x slower | Shared ractor for all lease maps |
Ractor::ActorHash |
400x slower | 280x slower | Additional ractor per map |
The map implementations namespaced under Ractor are from the ractor-sharing gem.
Also note that Ratomic::Map has a significant performance benefit over all other implementations when repeatedly writing to different keys in very large maps concurrently on a very high number of Ractors due to the underlying DashMap implementing data sharding. This benefit does not materializes if different Ractors share the keys they use, so its usefulness is slightly hampered by the fact that you cannot iterate over Ratomic's maps at all.
Counter performance
Caution
If counter performance is your application's bottleneck, Ruby might not be the right choice for you.
CRuby (4.0)
| Implementation | Increment | Read value | Integer size | Note |
|---|---|---|---|---|
| farce | fastest | fastest | 64-bit | |
| concurrent-ruby-ext | 1.2x slower | same-ish | 32-bit | no ractor support |
| ratomic | 1.5x slower | 1.6x slower | 64-bit | no overflow protection |
| ractor_safe | 2.1x slower | 2.1x slower | 64-bit | |
| concurrent-ruby | 12x slower | 5x slower | 32-bit | no ractor support |
Performance differences between implementations are consistent between single-threaded and multi-threaded benchmarks.
JRuby
| Implementation | Increment | Read value | Integer size |
|---|---|---|---|
| concurrent-ruby-ext | fastest | fastest | 64-bit |
| farce | 1.7x slower | 2.2x slower | 64-bit |
| concurrent-ruby | 15x slower | 9x slower | 32-bit |
Farce uses a JVM-specific Ruby implementation, concurrent-ruby uses the same Mutex-based implementation as on other platforms, and concurrent-ruby-ext comes with a Java implementation of an atomic counter (hence also the difference in integer size). The overhead in Farce can largely be attributed to Ruby dispatch overhead.
TruffleRuby
| Implementation | Increment | Read value | Integer size |
|---|---|---|---|
| farce | fastest | fastest | 64-bit |
| concurrent-ruby | 1.5x slower | 3.5x slower | 32-bit |
TruffleRuby's performance numbers are not as reliable as other Ruby implementations, and may vary significantly between runs, versions, and whether the GraalVM is in use and has warmed up. Neither farce nor concurrent-ruby use a counter written in C, so they should both be fully optimizable by the GraalVM.
Lock performance
Caution
If lock performance is your application's bottleneck, you might want to look into different data structures. Farce and concurrent-ruby provide plenty of options.
CRuby
Farce's locks are between 5% and 10% slower than Mutex for uncontended locks, and stay below a 50% performance penalty for highly contended locks between threads.
JRuby and TruffleRuby
Farce's locks have identical performance to Mutex (as they are a subclass of Mutex).
Queue performance
In the producer/consumer workload in benchmark/queue.rb, Ruby's built-in Thread::Queue and Thread::SizedQueue remain the fastest, but they do not support ractors.
| Implementation | Performance | Notes |
|---|---|---|
Thread::Queue |
Fastest | no ractor support |
Thread::SizedQueue |
1.4x slower | no ractor support |
Farce::Strict::Queue |
1.5x slower | only allows sharable objects |
Farce::Queue |
2.1x slower | |
RactorSafe::Queue |
2.9x slower | only allows sharable objects |
Ratomic::Queue |
9.8x slower | breaks ractor isolation |
RactorQueue |
30x slower | breaks ractor isolation |
Ractor::Port (multiplexing) |
100x slower |
Priority queue performance
Many gems implement a priority queue or comparable data structure. Farce's implementation is the only one that allows cross-ractor communication. The below numbers compare non-blocking APIs, as only Farce implements a blocking API as well.
| CRuby | JRuby | ||||
|---|---|---|---|---|---|
| ↓ Gem | Insertion order → | Random | Descending | Random | Descending |
| farce 0.1.0 | fastest | fastest | fastest | fastest | |
| rbtree 0.4.7 | 1.1x slower | 1.2x slower | – | – | |
| priority_queue_cxx 0.3.7 | 1.7x slower | 1.7x slower | – | – | |
| io-event 1.21.1 | 2.6x slower | 3.9x slower | 1.3x slower | 1.7x slower | |
| pqueue 2.2.0 | 9.5x slower | 8.0x slower | 2.2x slower | 1.9x slower | |
| lazy_priority_queue 0.1.1 | 9.9x slower | 7.9x slower | 3.9x slower | 3.2x slower | |
| philiprehberger-priority_queue 0.5.0 | 13x slower | 19x slower | 7.6x slower | 11x slower | |