Written by James Britt
Picture a familiar Rails support problem: a customer has requested a CSV export, the application has accepted it, and the download never arrives. You open the dashboard. Everything looks fine.
Somewhere behind that working interface, an export job is waiting for a worker that isn’t coming.
Explaining this with a green “All systems operational” badge makes an irritating problem worse. For a Ruby application, the public status page needs to account for what happens after a controller returns a response. Otherwise, it can tell a very different story from the one customers are experiencing.
Choose components that make sense outside your development team
Internally, you might investigate Puma processes, database connections, or a particular job class. Those details help you find a fault. They don’t necessarily help a customer decide whether to carry on working.
Organize the public page around services like the API, account access, report exports, and outgoing webhooks. Use names people already encounter in the application.
For a Rails reporting product, the distinction could look like this:
| Public component | What the Ruby team investigates |
|---|---|
| Dashboard and sign-in | Request errors, response times, and the login workflow |
| API | Failures and latency on representative endpoints |
| Report exports | Waiting jobs, failures, and completed files |
| File uploads | Upload and retrieval through the storage service |
| Outgoing webhooks | Delivery delays, failed attempts, and retry backlogs |
Your page may need fewer components. There’s little benefit in listing exports if your application doesn’t offer them. But if scheduled reports are the reason customers pay for the product, give that workflow its own status.
Know the limits of the Rails health endpoint
Rails includes a health endpoint, usually at /up. It reports whether the application booted without exceptions. It doesn’t verify that the database, Redis, or all the other services your application needs are working.
The route belongs in config/routes.rb:
get "up" => "rails/health#show", as: :rails_health_check
Check before adding it: newly generated Rails applications already include a health route.
Treat its response as one piece of evidence. For a fuller picture, ask whether a test account can sign in, whether an API request returns the expected record, and whether an export finishes within the time your team considers acceptable.
Use dedicated test data and limit the work each probe performs. A monitoring request shouldn’t place a real order or repeatedly email a customer.
Be careful with restart policies, too. A payment provider becoming unavailable is no reason to restart every Rails instance. The Rails documentation warns that dependency checks can produce this kind of unwanted restart behavior.
Watch what happens after perform_later
Take this line from a hypothetical Rails reporting application:
ReportExportJob.perform_later(report.id)
It requests background execution. It doesn’t tell you that a report now exists. Active Job supplies the interface, while a backend such as Solid Queue or Sidekiq handles execution.
Suppose the application accepts requests all afternoon, but no worker is consuming the exports queue. Checking the homepage won’t reveal the growing wait.
For exports, track how long eligible jobs sit in the queue, whether they complete, and how often they fail. Leave intentionally scheduled future jobs out of your overdue count. A worker heartbeat helps, but a running worker can still be listening to the wrong queue.
Customers don’t need that whole diagnosis. “Report exports are delayed” gives them the useful part.
Don’t close the incident just because a worker has restarted. Check the older requests as well as new submissions. If yesterday’s export is still stuck, its owner won’t consider the problem resolved.
Make sure the status page survives the outage
Adding a StatusController to your main application is tempting. You already have Rails, authentication, and somewhere to store incident records.
Then the database goes down, and your explanation goes with it.
Before choosing that arrangement, trace its dependencies. Can you publish an update without logging into the affected application? Will a failed deployment remove both the product and its status page? Does notification delivery depend on the same workers you’re investigating?
Keep the public page and its publishing mechanism independent of the application’s critical dependencies. Giving it a separate hostname won’t accomplish that by itself.
Tell people what they can still do
During an incident, developers might be inspecting a connection pool or comparing worker deployment revisions. A customer is wondering whether they can finish today’s work.
An opening update for an imagined export incident might read:
Report exports are taking longer than usual. Existing reports remain available. We are investigating the export workers and will post another update by 14:30 UTC.
Publish those details only after confirming them. The message identifies an affected feature, tells people what remains usable, and sets a time to check back.
Someone needs to own that next update. If the cause is still unknown at 14:30, say so. Support staff shouldn’t have to guess whether an unchanged message means the investigation has stopped.
Keep credentials, exception traces, customer identifiers, and internal hostnames in restricted incident records. Public updates should explain the impact without exposing those details.
When service recovers, say whether delayed work has cleared and whether customers need to retry. Check for duplicate-request risks before recommending that everyone submit again.
Give social posts somewhere permanent to point
Some users will look for your Ruby project’s maintainers on social media. A short post can help them find an update.
Keeping the only incident record there is another matter. Social platforms have their own outages, and a thread is awkward to follow when you arrive halfway through an investigation.
Maintain a stable incident address. Link to it from support replies, social posts, and application help pages. Where possible, include it on error pages that can still be served while Rails is unavailable.
That dated history also helps customers using your Ruby API compare your incident window with failures in their own logs.
Decide how much status infrastructure you want to maintain
A separate Rails status application may start with a couple of models and an incident form. Subscriptions, notification failures, publishing permissions, and the availability of that application all add work.
For a small Ruby team, a hosted service may be a practical choice. DevHelm offers uptime monitoring and public status pages with components and custom domains. It lists subscriber notifications as a paid-tier feature. Those are useful capabilities to evaluate against the needs of your Rails service.
Before allowing monitoring results to change public status, agree on what warrants an incident. A single failed probe might need investigation before publication. Conversely, a passing homepage check shouldn’t clear a confirmed problem with exports.
Keep a manual way to publish updates. Whatever detects the failure, somebody still has to explain it.
Keep a record that means something
Planned Rails maintenance deserves the same care. For a database migration or worker deployment, announce the expected customer impact and the maintenance window, with an explicit time zone.
After an outage, record which features were affected and for how long. If you publish uptime figures, explain their scope. Successful homepage checks cannot demonstrate that every queued report was available.
Start with a few important workflows and make their status accurate. A customer waiting for a report needs to know whether it is delayed, whether the team is working on it, and when to look for another update. Your Rails status page should answer those questions.
