Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
18 changes: 18 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,23 @@
# Changelog

## Unreleased

- `timeout` now covers every phase of the default transport, not only connect
and read. A stalled TLS handshake or a stalled upload used to fall back to
`Net::HTTP`'s own 60-second write timeout, well past both `timeout` and
`total_timeout`.
- `Net::WriteTimeout` is classified as a timeout, so `retry_timeouts` governs
it like the other two instead of it surfacing as an unretried
`TransportError`.
- `total_timeout` is enforced as a deadline. Each attempt is given the smaller
of `timeout` and the remaining budget, an attempt that would start with no
budget left raises `TimeoutError`, and a delay landing exactly on the
deadline now stops the retry loop rather than allowing one more attempt.
- The transport contract accepts an optional `timeout:` keyword carrying that
per-attempt budget. Transports that do not declare it are called unchanged.
- `timeout:` is validated when the client is built: nil, or a finite positive
number. Anything else raises `ConfigurationError`.

## 0.1.0 - 2026-09-18

Provider-neutral release. One `Client`, two providers behind it.
Expand Down
29 changes: 19 additions & 10 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -76,7 +76,7 @@ RubyDecisionModel::Client.new(
api_key: nil, # overrides the provider's env var
model: nil, # nil means the provider default; see aliases below
base_url: nil, # overrides the provider base URL
timeout: 5, # open and read timeout in seconds
timeout: 5, # per-attempt timeout in seconds, nil for none
retry: { max_retries: 2 }, # RetryPolicy or a Hash of overrides
transport: nil # see Transport
)
Expand Down Expand Up @@ -147,9 +147,11 @@ Retry behaviour follows the official Typesafe SDKs and lives in
| `retry_timeouts` | `true` | Retry open and read timeouts |
| `total_timeout` | `30.0` | Budget in seconds across attempts and delays; `nil` disables |

When the next delay would push past `total_timeout`, the client stops and
raises the last error instead of sleeping. The budget governs whether another
attempt starts; an attempt already in flight still runs to its own `timeout`.
`total_timeout` is a deadline, not just a gate between attempts. Each attempt
is given the smaller of `timeout` and what is left of the budget, so a single
slow attempt cannot outlive the whole call. When the next delay would reach or
pass the budget the client stops and raises the last error instead of sleeping,
and an attempt that would start with nothing left raises `TimeoutError`.

Invalid settings (a negative duration, a non-integer `max_retries`, a jitter
outside 0..1, a NaN budget) raise `ConfigurationError` when the client is built.
Expand All @@ -161,17 +163,24 @@ RubyDecisionModel::Client.new(retry: RubyDecisionModel::RetryPolicy.new(max_retr

## Transport

The client uses `Net::HTTP` by default. Inject `transport:` with any callable
that accepts `url:`, `headers:`, `body:` and returns
`[status, body_string, headers_hash]`. A two-element `[status, body_string]`
return is still accepted and treated as having no headers, which means no
`Retry-After` support and a nil `request_id`.
The client uses `Net::HTTP` by default, with `timeout` applied to all four of
its phases: connect, TLS handshake, write, and read.

Inject `transport:` with any callable that accepts `url:`, `headers:`, `body:`
and returns `[status, body_string, headers_hash]`. A two-element
`[status, body_string]` return is still accepted and treated as having no
headers, which means no `Retry-After` support and a nil `request_id`.

A transport that also declares a `timeout:` keyword (or `**`) is handed the
number of seconds this attempt may take, already clamped to what is left of
`total_timeout`. Transports that do not declare it are called exactly as
before.

## Errors

| Error | Meaning |
| --- | --- |
| `ConfigurationError` | No provider could be resolved, missing api_key, unknown provider, or bad `retry:` value |
| `ConfigurationError` | No provider could be resolved, missing api_key, unknown provider, or bad `timeout:` or `retry:` value |
| `RequestError` | Questions hash was empty |
| `TransportError` (`TimeoutError`) | Network or timeout failure after retries, carries `#cause_error` |
| `ApiError` | Non-2xx response, carries `#status`, `#body`, and `#headers` |
Expand Down
72 changes: 66 additions & 6 deletions lib/ruby_decision_model/client.rb
Original file line number Diff line number Diff line change
Expand Up @@ -41,7 +41,7 @@ def initialize(provider: nil, api_key: nil, model: nil, base_url: nil, timeout:
end

@model = @provider.resolve_model(model)
@timeout = timeout
@timeout = validate_timeout(timeout)
@transport = transport || default_transport
@sleeper = sleeper
@retry_policy = RetryPolicy.from(binding.local_variable_get(:retry))
Expand Down Expand Up @@ -89,12 +89,18 @@ def resolve_provider(provider, api_key:, base_url:)
end

def default_transport
lambda do |url:, headers:, body:|
lambda do |url:, headers:, body:, timeout: @timeout|
uri = URI.parse(url)
http = Net::HTTP.new(uri.host, uri.port)
http.use_ssl = uri.scheme == "https"
http.open_timeout = @timeout
http.read_timeout = @timeout
# All four, not just open and read: a stalled TLS handshake or a
# stalled upload otherwise falls back to Net::HTTP's own defaults
# (60s for the write), which blows through both `timeout` and the
# retry policy's total budget.
http.open_timeout = timeout
http.ssl_timeout = timeout
http.read_timeout = timeout
http.write_timeout = timeout

request = Net::HTTP::Post.new(uri.request_uri)
headers.each { |k, v| request[k] = v }
Expand All @@ -105,6 +111,34 @@ def default_transport
end
end

def validate_timeout(timeout)
return timeout if timeout.nil?
return timeout if timeout.is_a?(Numeric) && timeout.finite? && timeout.positive?

raise ConfigurationError, "timeout must be nil or a finite positive number, got #{timeout.inspect}"
end

# The transport contract grew a `timeout:` keyword so the client can hand
# each attempt what is left of the retry budget. Transports written
# against the old three-keyword contract still work: they are called the
# way they always were.
def transport_accepts_timeout?
return @transport_accepts_timeout unless @transport_accepts_timeout.nil?

parameters = @transport.respond_to?(:parameters) ? @transport.parameters : @transport.method(:call).parameters
@transport_accepts_timeout = parameters.any? do |kind, name|
kind == :keyrest || (%i[key keyreq].include?(kind) && name == :timeout)
end
end

def call_transport(url:, headers:, body:, timeout:)
if transport_accepts_timeout?
@transport.call(url: url, headers: headers, body: body, timeout: timeout)
else
@transport.call(url: url, headers: headers, body: body)
end
end

def perform_with_retry(url:, headers:, body:)
policy = @retry_policy
started_at = @clock.call
Expand All @@ -113,7 +147,8 @@ def perform_with_retry(url:, headers:, body:)
loop do
begin
status, response_body, response_headers = normalize_transport_result(
@transport.call(url: url, headers: headers, body: body)
call_transport(url: url, headers: headers, body: body,
timeout: attempt_timeout(policy, started_at))
)
rescue Error
raise
Expand Down Expand Up @@ -146,7 +181,32 @@ def perform_with_retry(url:, headers:, body:)
def budget_exceeded?(policy, started_at, delay)
return false if policy.total_timeout.nil?

(@clock.call - started_at) + delay > policy.total_timeout
# >=, not >: a delay that lands exactly on the deadline has used the
# whole budget, and the attempt after it would start with nothing left.
(@clock.call - started_at) + delay >= policy.total_timeout
end

# What is left of the budget, or nil when there is no budget. Handed to
# the transport so a single attempt cannot outlive the whole call: with
# `timeout: 5` and 2s of budget left, the attempt gets 2s.
def remaining_budget(policy, started_at)
return nil if policy.total_timeout.nil?

policy.total_timeout - (@clock.call - started_at)
end

def attempt_timeout(policy, started_at)
remaining = remaining_budget(policy, started_at)
return @timeout if remaining.nil?

if remaining <= 0
raise TimeoutError.new(
"request budget of #{policy.total_timeout}s was exhausted before the attempt started",
cause_error: nil
)
end

@timeout.nil? ? remaining : [@timeout, remaining].min
end

def normalize_transport_result(result)
Expand Down
5 changes: 4 additions & 1 deletion lib/ruby_decision_model/retry_policy.rb
Original file line number Diff line number Diff line change
Expand Up @@ -21,7 +21,10 @@ class RetryPolicy < Data.define(
)
DEFAULT_STATUSES = ([408, 429] + (500..599).to_a).freeze

TIMEOUT_EXCEPTIONS = [Net::OpenTimeout, Net::ReadTimeout].freeze
# Net::WriteTimeout fires when the request body stalls on the way out.
# It is as much a timeout as the other two, and leaving it off this list
# made it a plain TransportError that was never retried.
TIMEOUT_EXCEPTIONS = [Net::OpenTimeout, Net::ReadTimeout, Net::WriteTimeout].freeze
CONNECTION_EXCEPTIONS = [
Errno::ECONNRESET,
Errno::ECONNREFUSED,
Expand Down
Loading