Benchmarks

Farce also aims to be fast and efficient, aiming for anywhere between a minimal overhead to outperforming other options.

Any numbers quoted here are to be taken with a grain of salt:

  • They are based on micro-benchmarks, which may not reflect real-world performance.
  • They are momentary snapshots. Gems and Ruby implementations are constantly evolving, so these numbers may not be accurate in the future.
  • Concrete numbers are largely measured on a local machine, which may not reflect your deployment environment.

You should measure the performance of your own application under realistic conditions.

Map performance

A shared map is a very common data structure for tracking state. While the read and write performance is hopefully not the bottleneck for your application.

Map implementation Read Write Notes
Hash fastest fastest Not ractor-shareable (when mutable), not thread-safe
Concurrent::Hash 1.1x slower 1.1x slower Not ractor-shareable
Farce::Strict::Map 1.2x slower 1.1x slower
Farce::Unshared::Map 1.2x slower 1.1x slower Not ractor-shareable
Farce::Map 1.3x slower 1.1x slower
Concurrent::Map 1.4x slower 3.2x slower Not ractor-shareable
Hash + Mutex 2.9x slower 2.9x slower Not ractor-shareable
Ractor::LockHash 2.9x slower 3.4x slower Not fiber-friendly
Ratomic::Map 3.1x slower 2.7x slower Breaks isolation, not fiber-friendly
Ractor::KeyLockHash 3.2x slower 2.5x slower Not fiber-friendly
Farce::LRUMap 3.4x slower 7.6x slower Automatic eviction
Farce::LFUMap 3.5x slower 7.6x slower Automatic eviction
Farce::Unsafe::TreeMap 3.7x slower 4.1x slower Ordered entries, not thread-safe, not ractor-shareable
HashWithIndifferentAccess 4.5x slower 5.3x slower Not ractor-shareable
RactorSafe::HashMap 4.8x slower 4.8x slower
Farce::TreeMap 5.5x slower 40x slower Ordered entries
Farce::LeaseMap 30x slower 40x slower Shared ractor for all lease maps
Ractor::ActorHash 400x slower 280x slower Additional ractor per map

The map implementations namespaced under Ractor are from the ractor-sharing gem.

Also note that Ratomic::Map has a significant performance benefit over all other implementations when repeatedly writing to different keys in very large maps concurrently on a very high number of Ractors due to the underlying DashMap implementing data sharding. This benefit does not materializes if different Ractors share the keys they use, so its usefulness is slightly hampered by the fact that you cannot iterate over Ratomic's maps at all.

Counter performance

Caution

If counter performance is your application's bottleneck, Ruby might not be the right choice for you.

CRuby (4.0)

Implementation Increment Read value Integer size Note
farce fastest fastest 64-bit
concurrent-ruby-ext 1.2x slower same-ish 32-bit no ractor support
ratomic 1.5x slower 1.6x slower 64-bit no overflow protection
ractor_safe 2.1x slower 2.1x slower 64-bit
concurrent-ruby 12x slower 5x slower 32-bit no ractor support

Performance differences between implementations are consistent between single-threaded and multi-threaded benchmarks.

JRuby

Implementation Increment Read value Integer size
concurrent-ruby-ext fastest fastest 64-bit
farce 1.7x slower 2.2x slower 64-bit
concurrent-ruby 15x slower 9x slower 32-bit

Farce uses a JVM-specific Ruby implementation, concurrent-ruby uses the same Mutex-based implementation as on other platforms, and concurrent-ruby-ext comes with a Java implementation of an atomic counter (hence also the difference in integer size). The overhead in Farce can largely be attributed to Ruby dispatch overhead.

TruffleRuby

Implementation Increment Read value Integer size
farce fastest fastest 64-bit
concurrent-ruby 1.5x slower 3.5x slower 32-bit

TruffleRuby's performance numbers are not as reliable as other Ruby implementations, and may vary significantly between runs, versions, and whether the GraalVM is in use and has warmed up. Neither farce nor concurrent-ruby use a counter written in C, so they should both be fully optimizable by the GraalVM.

Lock performance

Caution

If lock performance is your application's bottleneck, you might want to look into different data structures. Farce and concurrent-ruby provide plenty of options.

CRuby

Farce's locks are between 5% and 10% slower than Mutex for uncontended locks, and stay below a 50% performance penalty for highly contended locks between threads.

JRuby and TruffleRuby

Farce's locks have identical performance to Mutex (as they are a subclass of Mutex).

Queue performance

In the producer/consumer workload in benchmark/queue.rb, Ruby's built-in Thread::Queue and Thread::SizedQueue remain the fastest, but they do not support ractors.

Implementation Performance Notes
Thread::Queue Fastest no ractor support
Thread::SizedQueue 1.4x slower no ractor support
Farce::Strict::Queue 1.5x slower only allows sharable objects
Farce::Queue 2.1x slower
RactorSafe::Queue 2.9x slower only allows sharable objects
Ratomic::Queue 9.8x slower breaks ractor isolation
RactorQueue 30x slower breaks ractor isolation
Ractor::Port (multiplexing) 100x slower

Priority queue performance

Many gems implement a priority queue or comparable data structure. Farce's implementation is the only one that allows cross-ractor communication. The below numbers compare non-blocking APIs, as only Farce implements a blocking API as well.

CRuby JRuby
↓ Gem Insertion order → Random Descending Random Descending
farce 0.1.0 fastest fastest fastest fastest
rbtree 0.4.7 1.1x slower 1.2x slower – –
priority_queue_cxx 0.3.7 1.7x slower 1.7x slower – –
io-event 1.21.1 2.6x slower 3.9x slower 1.3x slower 1.7x slower
pqueue 2.2.0 9.5x slower 8.0x slower 2.2x slower 1.9x slower
lazy_priority_queue 0.1.1 9.9x slower 7.9x slower 3.9x slower 3.2x slower
philiprehberger-priority_queue 0.5.0 13x slower 19x slower 7.6x slower 11x slower