Written by James Britt
A Ruby web crawler may start out with three proxy servers and a call to Array#sample. That can work at first. Later, a proxy begins to fail, retries repeat more work than expected, and the database fills with pages that contain no useful records.
Your Ruby code still has to determine which server to use as its proxy, whether another attempt is justified, and what constitutes good data.
A simple implementation is easy to follow. The examples below implement an HTTPS GET fetcher using Net::HTTP, a round-robin selector and a bounded retry loop.
Put Proxy Settings in One Place
When establishing a connection, Net::HTTP allows you to specify a proxy host, port, and optionally credentials. Those docs cover Ruby 2.6. Check the docs for the Ruby and net-http versions your app uses.
Faraday accepts proxy settings at a connection level and routes the requests through an adapter. HTTParty provides http_proxy on the client class and proxy options on each request individually.
Separate proxy selection from extracting product names. When you switch proxy endpoints you should not have to change the code that parses documents for product names.
Create a method that performs one request:
require "net/http"
require "uri"
require "thread"
def get_via_proxy(url, proxy)
target = URI(url)
unless target.is_a?(URI::HTTPS) && target.hostname
raise ArgumentError, "Expected an HTTPS URL"
end
http = Net::HTTP.new(
target.hostname, target.port,
proxy.fetch(:host), proxy.fetch(:port),
proxy[:user], proxy[:password]
)
http.use_ssl = true
http.open_timeout = 5
http.read_timeout = 15
http.write_timeout = 15
http.max_retries = 0
request = Net::HTTP::Get.new(target.request_uri)
http.start { |connection| connection.request(request) }
end
This uses an HTTP proxy to reach an HTTPS destination, keeps certificate verification enabled and opens a new connection per request. Load the credentials from environment variables or a secure store.
A read timeout is not a total download deadline for the page. Net::HTTP applies it to individual reads. A slow response can continue sending data for longer. A production fetcher may also need a total time budget and a response-size limit.
Setting max_retries to 0 will prevent Net::HTTP’s automated retries from occurring under our retry loop.
Use a Predictable Method for Selecting Proxies
Using random choice can select the same proxy twice in a row. Random selection does not stop you from gathering metrics. Round-robin selection gives each proxy a turn while all entries remain eligible.
Below is a selector that rejects an empty array and protects the index with a Mutex:
class ProxyRotator
def initialize(proxies)
raise ArgumentError, "Proxy list is empty" if proxies.empty?
@proxies = proxies.map { |proxy| proxy.dup.freeze }.freeze
@index = 0
@mutex = Mutex.new
end
def next_proxy
@mutex.synchronize do
proxy = @proxies.fetch(@index)
@index = (@index + 1) % @proxies.length
proxy
end
end
end
The Mutex only protects proxy selection. The network request occurs after the Mutex has been released. Thus, a single slow network connection cannot block every thread attempting to select a proxy.
This class copies and freezes the configuration hashes. Their values are not deeply frozen, so treat these settings as immutable after initialization.
The selector has a small job. It does not cap concurrent requests, enforce destination rate limits or remove unhealthy endpoints. Equal selection counts do not imply equal bandwidth or completion times.
The Mutex synchronizes threads accessing this object in one Ruby process. Separate Sidekiq processes or containers need external coordination if selections or limits must be shared.
Create a Single Loop to Manage the Retry Budget
It is easier to see how many connection attempts were made with an explicit loop. An exception handler retry can cause more of a method to run than you intend.
This version retries selected transport failures on GET requests:
NETWORK_ERRORS = [
Net::OpenTimeout, Net::ReadTimeout, Net::WriteTimeout,
SocketError, EOFError,
Errno::ECONNREFUSED, Errno::ECONNRESET
].freeze
def fetch(url, rotator, attempts: 3)
unless attempts.is_a?(Integer) && attempts.positive?
raise ArgumentError, "attempts must be a positive integer"
end
attempts.times do |index|
proxy = rotator.next_proxy
begin
return get_via_proxy(url, proxy)
rescue *NETWORK_ERRORS
raise if index == attempts - 1
sleep(0.25 * (2**index))
end
end
end
If attempts is set to 3 then the method will call get_via_proxy at most 3 times. If all 3 attempts fail then the third exception will be propagated back to the caller. Each retry selects the next configured proxy.
Destination responses, including 403, 429 and 503, are returned without another attempt. The caller must categorize them. A proxy authentication failure while establishing the HTTPS tunnel can raise an exception before the destination request is sent. This loop does not catch that failure or TLS certificate errors.
You should not copy this retry strategy onto a POST that generates orders or charges accounts. A missing response does not mean the server performed no work.
To tie everything together, provide two sets of proxy configs via environment variables:
proxies = (1..2).map do |number|
prefix = "PROXY_#{number}"
{
id: prefix,
host: ENV.fetch("#{prefix}_HOST"),
port: Integer(ENV.fetch("#{prefix}_PORT")),
user: ENV.fetch("#{prefix}_USER"),
password: ENV.fetch("#{prefix}_PASSWORD")
}
end
rotator = ProxyRotator.new(proxies)
response = fetch("https://example.com/data", rotator)
puts response.code
The destination is a placeholder. Replace it with an endpoint that you have authorization to query and supply your proxy settings. This small example does not follow redirects, persist cookies or pool connections.
Verify the Response Prior to Storing It
A response with status 200 can still be invalid because it contains incorrect content. The parser should identify the expected page structure or payload before storing records in the database.
For JSON, parse the response body and check required keys and value types. For HTML, search for expected page structures. A broad search for ‘captcha’ can mistake an article about CAPTCHAs for a challenge page.
Store response categories separately. A 403 signifies the server denied the request; a 407 signifies proxy authentication was needed; a 429 signifies you reached rate limits; and a 503 signifies the service was unavailable. None of these alone proves that a specific proxy is defective.
For responses with 429 status codes, honor the Retry-After header when supplied and enforce destination-specific rate limits. Switching between proxies does not relieve you of enforcing rate limits.
Log a proxy id, destination hostname, duration, status, and validation results. Avoid logging credentials, session cookies or entire response bodies by default.
Keep a Workflow on the Same Proxy
Some workflows require consistent routing throughout multiple requests. Assign a proxy when the workflow starts, and keep that assignment until it completes or fails.
Keeping the same proxy does not preserve cookies by itself. Workflows require their own cookie jars and session state. A fixed gateway hostname does not guarantee a fixed exit IP address, either. That depends on the provider’s session behavior.
Do not expire an assignment mid-way through filling out a form or completing a series of pages. If assignment expiration is necessary then use a monotonic clock to measure elapsed time. Remove finished or abandoned items so that your session maps remain bounded.
The fetch method described above rotates after transport failures. A sticky workflow needs a separate policy that keeps its assigned proxy and decides whether to restart the whole workflow.
Track Successes and Failures Without Changing History
A pool that tracks health needs separate counters for attempts, completed attempts, successes and failures. Record one outcome per completed attempt. Keep the failure streak used for cooldown decisions in a separate field.
Assume two consecutive failures and one subsequent success for a given proxy. The historical failure counter still records two failures. The failure streak resets to zero after the success.
Decreasing the failure counter on success would understate the failure rate.
Calculate failure rates as failures divided by completed attempts within a specified timeframe where failed attempts are included in the denominator. Track transport issues separately from destination responses and parser errors. A changed HTML template should not cause every proxy to be marked unhealthy.
Cooldowns require a recovery path. Schedule a limited probe after the cooldown ends, and decide what the caller should do if every proxy endpoint is unavailable. Otherwise, an exclusion rule can starve the pool forever.
Store Provider-Specific Data in Configurations
Managed gateways can move exit address selection outside your application. However, Ruby still needs timeouts, session constraints, response validation logic and meaningful logs.
The original NetNut rotating proxies link is retained here as a historical reference. As of September 2026, it returns a domain seizure notification. Also, Google reported action against the network in July 2026. The old gateway settings should not be treated as a verified current configuration.
Read the active provider’s endpoint and session parameters from configuration. This enables changing suppliers without modifying your parsers or retry policies.
Keep three decisions visible in your Ruby application: which route to use, how to issue the request and whether the response belongs in the dataset. When a job fails, those boundaries give you somewhere specific to start looking.
