
TL;DR
Rails is slow, but Ruby can be fast. A lean Ruby stack works very well, 10x more throughput than Rails in my own benchmarks, and LLMs can help maintain it for you.
What’s up?
A lot of discussion and, sometimes heated, debate arose from DHH’s latest Rails World keynote presentation. His new found love for both LLM generated code and fast, efficient runtime seem to come as a surprise to many. Especially when combined with what “seems” to be a ditching of Ruby on Rails at 37Signals. Many are applauding the bravery, many are questioning the audacity of ditching Rails in the Rails World keynote and some are providing a different perspective. Sam Ruby’s Roundhouse and Matz’s Spinel together are offering an argument for Rails in this new, agentic coding world. I want to offer my own perspective based on my past and current experience writing both Rails and vanilla Ruby applications.
Rails is dog slow
I think this is the core issue in all this. Rails is dog slow.
When first released, it was a breath of fresh air in web development simplicity and speed of delivery. This has changed a bit lately as the framework matured, but it is still, LLMs aside for now, one of the best options for a single person or a small team to quickly build a web application or an app back end.
But DHH was steering Rails, and for most of his involvement with it he was focused on developer productivity over most other things, including performance. Indeed, developer time, in many contexts, including DHH’s direct context, was a lot more expensive than any infra costs incurred by slower runtimes. Just throw hardware, which is getting cheaper by the day, at the problem and be happy.
As a result of this, there were these takes in the Rails community. Like when Rails developers seem to accept provisioning one core per one to two requests per second when sizing machines! Other communities were perplexed at such takes, and rightly so, but that didn’t affect the community much. At least DHH was showing benchmarks with 10 req/s per core on very modern hardware, still abysmal but not as absurd as 1 req/s.
With the release of the Solid set of gems, Solid Cache, Queue and Cable, default Rails performance took to new levels of slowness. See, these ActiveRecord based abstractions are really nice for something like a small Once based deployment of Campfire and similar applications. But it was sold, as usual, as the new way to do stuff and most new deployments use them by default, being, the “Rails way”!
The new reality
After LLMs became more than half decent in producing code that actually works, people realized they can easily port their applications to faster runtimes/frameworks. Potentially losing the ability to read the code in the process but not many would care really. All of a sudden DHH realized the “Rails way” is super slow, and that the future is Rusty, you get performance and safety and everything you advocated against for years because you prioritized productivity. Now with LLMs and unlimited tokens, you are no longer tied to what you yourself can do in your and your team’s limited time. But to what the LLMs can produce and if they can implement lower level code that runs faster and saves you infra costs then why not?
Not to mention that hardware is no longer getting cheaper by the day, again, after LLMs demand for memory and compute and power has skyrocketed. Resulting in higher prices for everyone, especially for RAM. Throwing hardware at the problem might work still, but it looks a lot less attractive when you compare the cost per request back and then.
So now suddenly the realization that Rails is slow is sinking in and questions raised are like which role should it play in the development cycle in the future? After all, the port to Rust was partly successful because it started from a very well defined behavior of an existing Rails application. Should every Rails app out there be ported to Rust? Where should everyone go from here?
RoR fights back
I think one of the best takes I have seen so far was Sam Ruby’s both philosophical and practical responses to the current situation. Sam’s logic builds on the idea that a more abstracted approach results in much easier time for LLMs to fully comprehend and expand a codebase, thus Rails would be the smarter choice for LLMs to write code in. And then, to also cover the runtime efficiency, he bring in his Roundhouse tool that transpiles Rails applications into different targets, one of them is Spinel, Matz’s new AOT compiler for Ruby, resulting in much faster operation (between 2 to 4x slower than Rust currently) and much smaller runtime size, all while maintaining the Rails application as is which is awesome. I hope this direction matures and we come to see even better results from it. But what I want more is to see it expand beyond Rails and to other Ruby frameworks because that’s exactly what I want to discuss next.
Ruby is NOT dog slow
Actually, for web and IO applications in general, Ruby is quite decent. With the right set of tools it can offer very competitive performance while remaining a joy to write and maintain. I have been working with and nurturing my own home grown stack of years. It doesn’t have the polish or sophistication of something like Rails but I already have multiple applications deployed in production that rely on that home-made stack. LLMs made it even easier to maintain and improve those applications. In summary the stack is composed of:
- Rage-Iodine (the app server component) with a Fiber scheduler and WS support
- Extralite (for SQLite connectivity) + a mini ORM with (almost) zero runtime string manipulation
- Hanami::API for routing + a home grown controller framework + ERB for server side templates
- Application glue (startup, preloading, etc.)
- Flow.js for client side flow control of either JS templates (Sketch.js) + JSON or rendered HTML fragments
I have already been seeing great performance, and stability, from that stack and I have been using it, refining and expanding it to accommodate my use cases.
So I went to Opus 5.5 and asked it to port the Rails blog to my stack and to also port it to Kemal (Crystal). I even added a version (called rails-mode vs fast-mode for the original) that replicated the exact Rails security mechanisms (CSRF, encrypted sessions, HMAC, etc.) instead of my fast-mode choices, I then benchmarked the 4 implementations using the following setup:
- 7840HS CPU with 8 cores and 16 threads
- Ubuntu Linux running inside a VirtualBox VM
- Ruby 4.0.1 + YJIT for all Ruby apps
- Crystal 1.21 with the
--releasebuild flag - Rails 8.1.4, Puma 8.0.2, Redis 6.0.0 (gem), SQLite3 2.9.6 (gem), Extralite 3.1.1, Rage-Iodine 5.5
- Rails app had 16 puma workers with 1 thread each
- Ruby app had 16 Iodine workers with 1 thread each, in production mode
- Crystal app had a single worker process with 16 worker threads
HTTP Performance
wrk was used to send requests to different routes, 128 connections and 4 threads.
Throughput

All Ruby processes had YJIT enabled. Rails was using Puma, Kemal was used as is in a multi-threaded Crystal process. You can see that almost everything here is an order of magnitude faster than Rails.
Latency
| Route | Ruby stack, fast mode p50 / p99 (ms) | Ruby stack, Rails mode p50 / p99 (ms) | Crystal p50 / p99 (ms) | Rails p50 / p99 (ms) |
|---|---|---|---|---|
/articles.json | 0.95 / 5.13 | 1.16 / 7.66 | 1.54 / 10.02 | 18.37 / 132.79 |
/articles | 1.40 / 5.95 | 4.26 / 16.08 | 3.27 / 18.05 | 37.25 / 109.81 |
/articles/1.json | 0.95 / 4.96 | 1.08 / 5.32 | 1.16 / 9.07 | 11.58 / 23.57 |
/articles/1 | 1.41 / 6.80 | 4.27 / 11.84 | 1.72 / 31.32 | 35.80 / 80.00 |
The fast mode vanilla Ruby stack delivered the best latencies throughout. Rails delivered the worst numbers throughout, sometimes by an order of magnitude (again).
Note that this is a very simple app with a very lean request path, an app that does more work might see the gap shrink between Rails and the other Ruby implementations. In contrast to that statement, I believe Rails keeps adding overhead as you use more parts of it. ActiveJob adds overhead, the view cache adds overhead, many Rails gems add overhead, my personal experience is that the gap actually widens unless you are doing more direct computation/memory work like encryption or (de)compression.
WebSocket Broadcast Performance
The Crystal Cable implementation was in-process, in-memory as there was a single process with many threads. The Ruby stack implementation relies on Iodine’s inter-process pub-sub mechanism, orchestrated by the master process. The numbers shared here for Rails were using the Redis version, which, while still being slow, were a lot faster than the “solid” version.
A client using 8 iodine workers was created to connect to the application using web sockets and 500,000 deliveries were performed.
1,000 connected clients
| App | Throughput (deliveries/s) | p50 (ms) | p99 (ms) | max (ms) |
|---|---|---|---|---|
| Ruby stack, fast mode | 302,438 | 1.81 | 8.16 | 8.97 |
| Ruby stack, Rails mode | 196,903 | 2.33 | 16.82 | 37.86 |
| Crystal | 88,741 | 6.01 | 15.55 | 18.85 |
| Rails, Redis cable, async jobs | 30,692 | 19.89 | 41.47 | 54.35 |
10,000 connected clients
| App | Throughput (deliveries/s) | p50 (ms) | p99 (ms) | max (ms) |
|---|---|---|---|---|
| Ruby stack, fast mode | 471,950 | 12.09 | 22.51 | 26.57 |
| Ruby stack, Rails mode | 457,259 | 12.00 | 22.90 | 26.07 |
| Crystal | 101,924 | 60.76 | 134.90 | 145.17 |
| Rails, Redis cable, async jobs | 64,770 | 71.39 | 153.41 | 199.02 |
Again, an amazing showing for Ruby, it is doing very well, with help from Iodine, to deliver great performance on par or better than, fast, compiled runtimes. Rails is trailing with a wide margin again. The roughly 10x gap over Rails shows up again here.
Memory footprint
I checked the PSS for all applications after the HTTP benchmark and then again after the WebSocket benchmark. Here are the results:
| When measured | Ruby stack, fast mode (MB) | Ruby stack, Rails mode (MB) | Crystal (MB) | Rails (MB) |
|---|---|---|---|---|
| After the HTTP runs | 718.4 | 786.9 | 44.8 | 1,712 |
| After HTTP and WebSocket runs | 819 | 902 | 684.3 | 3,424 |
Naturally, the single process Crystal leads the pack here, but vanilla Ruby is more than double efficient RAM wise than Rails after the HTTP runs and over 4X more efficient after the WebSocket runs, it even comes within 20% of Crystal after the WebSocket run (I am not sure if there is something wrong with Kemal here though).
I expect that if I steer this framework to be Spinel compatible I could AOT compile my apps to a much smaller size than now, but I am not in a rush, the current numbers are still very good.
Your present is my past
I have been extracting such performance from Ruby since forever. Building really high speed, high scale applications while enjoying every bit of the development journey. I mentioned that before but one day a friend came to consult me because his 400K MAU Rails app was costing him thousands in AWS monthly bills. I didn’t know how to tell him back then that my 2M MAU “Ruby” apps (two combined), cost me $70/month to host. And they were complex beasts with a lot going on under the hood.
My thesis was that Ruby is awesome for the web, and you can, with little effort, extract a lot of performance from it, and I did. I welcome those that are discovering performance now though, it’s nice to see people realizing this, even if they took a really long path to it.
What about the future?
I have been more than enjoying working on my projects with smarter models. They write almost all the code now. But we plan together, identify bottlenecks, suggest experiments and validate guesses. The process is awesome, fast and productive. But I see my input being valuable both at the high level architecture decisions and the deep level optimization efforts.
Thanks to Ruby, the code base is well separated and is genuinely small and pleasant to deal with. I have been working with LLMS to not just produce the stack but to also enhance the docs targeted at LLMs quickly picking up and working with it for old and new projects.
Opus took only a couple minutes to port that miniature app, and it worked from the first run. I would love to also port the once-campfire app, but I don’t have the time nor do I have Unlimited TokensTM to throw at such experiments.
It’s about time
This is not an attempt to speculate about the future of Rails, it’s a discussion about the Ruby developer who does mostly Rails when faced with the idea that leaders in the community are pivoting to LLM generated Rust since the bottleneck has shifted from development time to runtime.
It’s about when and how you should react, and what to do next. Whether to optimize for development time or runtime, or both? While still using and enjoying Ruby. I present my own experience as one option and linking to Sam Ruby’s and Matz’s work as another. I am sure more will pop up.
In Conclusion
Ignore the drama, do things you love and embrace the tools of the future. You don’t have to use them to produce code you don’t want to look at or discuss.
Also pick the best tool for the job, in one case Ruby object sizes and GC overhead where too much for a lookup service with huge sets of in-memory data. It was decent but a port to Crystal freed up a lot of memory and more than doubled the performance.
Be curious, ask questions and check sources. This is an amazing opportunity to not just do more, but to also learn a-lot more in the process. As you move forward you will pickup the skill of knowing what you want to black-box and what you need to properly understand.
Life is not a race, but if we need a race analogy here then it would not be a sprint, but rather a marathon.
Leave a comment