---
title: 'Resilient integrations: timeouts, retries, circuit breakers, and idempotency | DevSense'
description: 'Laravel payment and API integrations without double charges or cascading failures: timeouts, backoff, circuit breakers, idempotency keys, reconciliation.'
faq:
    - { question: 'Why is a timeout the most dangerous outcome of an external API call?', answer: "With an error, you know the operation didn't happen; with a success, you know it did. With a timeout, you know nothing: the provider may have charged the card and simply failed to respond in time. A blind retry in this case leads to a double charge, so money operations are retried only with an idempotency key or after checking the status with the provider." }
    - { question: 'Which errors are safe to retry?', answer: "Connection errors, 429 responses, and 5xx responses, provided the operation is idempotent or protected by an idempotency key. Retrying 4xx responses (other than 408 and 429) is pointless: the request is invalid and will be rejected again. Retries use exponential backoff with random jitter so you don't hit the provider with a request storm." }
    - { question: 'How does a circuit breaker differ from a retry?', answer: 'A retry tries to ride out a brief failure of a single request. A circuit breaker protects your system from a prolonged provider outage: after a series of failures it stops sending requests and immediately returns an error or a fallback, without tying up workers on waiting. After a pause it lets a probe request through and, if that succeeds, reopens traffic.' }
    - { question: "How do you guarantee that a customer isn't charged twice?", answer: "You can't get an exactly-once guarantee for message delivery, but you can achieve exactly one effect. To do that, each operation gets a unique key that is saved in the database under a UNIQUE constraint before the provider is called, and is passed to the provider in the Idempotency-Key header. A repeated request or a repeated webhook with the same key returns the already stored result instead of creating a new operation." }
published: '2026-09-28'
---
# Resilient integrations: timeouts, retries, circuit breakers, and idempotency

The payment provider started responding in 30 seconds instead of 300 milliseconds. A minute later the whole site was down, including pages that had nothing to do with payments: every PHP-FPM worker was stuck waiting for a response. When the provider recovered, a second problem surfaced: some users had paid for their order twice. The code retried the request after a timeout, but the first request had actually gone through. Not a single line of code was "wrong." The external call was simply written as if it were a local function call.

**Related guides:** [Microservice patterns: Saga, CQRS, Circuit Breaker](../microservices/microservice-patterns) · [Distributed transactions](database-and-distributed-transactions) · [Message queues compared](message-queues-compared)

## Contents

* [The three outcomes of an external call](#three-outcomes)
* [Timeouts: a budget, not a default](#timeouts)
* [Retry: what, when, and how to retry](#retry)
* [Idempotency: one effect instead of two charges](#idempotency)
* [Incoming webhooks and concurrent debits](#webhooks)
* [Circuit Breaker: stop calling someone who doesn't answer](#circuit-breaker)
* [Isolation: a separate queue for each provider](#bulkheads)
* [Where this approach stops working](#limitations)
* [Common Mistakes](#common-mistakes)
* [Checklist](#checklist)
* [Self-Test Quiz](#self-test-quiz)

---

<a id="three-outcomes"></a>
## The three outcomes of an external call

A local method call has two outcomes: it returns a result or it throws an exception. A network call has three:

1. **Success**: the provider performed the operation and you received the response.
2. **Failure**: the provider responded that the operation was not performed (or the connection was never established).
3. **Unknown**: a timeout, a connection dropped after the request was sent, a 502 from a load balancer. The operation may or may not have been performed.

**A reliable integration is code that explicitly handles the third outcome: it limits how long it waits, retries only safe operations, stops calling a provider that isn't working, and uses idempotency keys to turn retries into a single effect.**

Everything else in this article is about not confusing "unknown" with "failure."

---

<a id="timeouts"></a>
## Timeouts: a budget, not a default

Laravel's HTTP client waits up to 30 seconds for a response by default. Let's work out what that means for a site with 50 PHP-FPM workers when the provider "hangs":

* each request to the checkout page ties up a worker for 30 seconds;
* at 2 requests per second, all 50 workers are busy after 25 seconds;
* the rest of the site starts returning 502, even though nothing is wrong with those pages.

A timeout is the budget you're willing to give the provider, based on its normal latency and on how many workers you can afford to risk:

```php
// app/Services/Payments/PaymentGatewayClient.php
<?php

declare(strict_types=1);

namespace App\Services\Payments;

use Illuminate\Http\Client\PendingRequest;
use Illuminate\Support\Facades\Http;

final class PaymentGatewayClient
{
    private function http(): PendingRequest
    {
        return Http::baseUrl(config('services.gateway.url'))
            ->withToken(config('services.gateway.token'))
            ->connectTimeout(2)  // establish the TCP/TLS connection
            ->timeout(5)         // receive the full response
            ->acceptJson();
    }
}
```

> [!TIP]
> **Move slow calls out of the user's HTTP request.** If the provider takes seconds to respond, send the request from a queued job and show the user a "payment is processing" status. That way a slow provider ties up queue workers rather than web workers.

---

<a id="retry"></a>
## Retry: what, when, and how to retry

A retry helps you ride out a brief failure: a pod restart on the provider's side, a network blip, a `429` response. But retries come with three conditions.

**1. Retry only what is safe to retry.** `GET` requests, status checks, operations with an idempotency key. A charge without a key must never be retried after a timeout.

**2. Retry only retryable errors.** Connection errors, `429`, `5xx`. A `422` or `400` response means the request is invalid, and a retry will get the same response.

**3. Retry with an increasing delay and random jitter.** If a thousand clients all retry exactly one second later, the provider takes a synchronized hit at the very moment it recovers.

```php
// app/Services/Payments/PaymentGatewayClient.php (continued)
use Illuminate\Http\Client\ConnectionException;
use Illuminate\Http\Client\RequestException;
use Throwable;

public function paymentStatus(string $paymentId): array
{
    return $this->http()
        ->retry(
            3,
            // Exponential backoff with jitter: ~200, ~400, ~800 ms
            fn (int $attempt): int => (int) (100 * 2 ** $attempt + random_int(0, 100)),
            fn (Throwable $e): bool => $e instanceof ConnectionException
                || ($e instanceof RequestException
                    && ($e->response->serverError() || $e->response->status() === 429)),
        )
        ->get("/payments/{$paymentId}")
        ->throw()
        ->json();
}
```

For queued jobs, the same thing is configured through job properties:

```php
// app/Jobs/CapturePayment.php
public int $tries = 5;

/** @return list<int> Delays between attempts, in seconds */
public function backoff(): array
{
    return [10, 30, 60, 300];
}
```

> [!WARNING]
> **Retries multiply.** If the HTTP client makes 3 attempts, the queued job makes 5, and the calling service retries too, the provider can receive 15+ requests for a single operation. Retry at one level only.

---

<a id="idempotency"></a>
## Idempotency: one effect instead of two charges

You can't get an exactly-once guarantee over a network: either you risk losing a message or you risk delivering it twice. The realistic goal is **at-least-once delivery plus idempotent processing**, which yields exactly one effect.

The flow for an outgoing charge:

1. Generate the operation key **before** the call and save the operation in the database with a `pending` status and a UNIQUE constraint on the key.
2. Pass the key to the provider in the `Idempotency-Key` header. Providers such as Stripe will return the result of the first request when they receive a repeat with the same key, instead of performing the operation again.
3. Save the result. On a timeout, set the status to `unknown` and find out what actually happened by querying the status, not by blindly retrying the charge.

```php
// database/migrations/2026_09_28_100000_create_payments_table.php
Schema::create('payments', function (Blueprint $table) {
    $table->id();
    $table->uuid('idempotency_key')->unique();
    $table->foreignId('order_id')->constrained();
    $table->unsignedBigInteger('amount');         // in minor units (kopecks/cents)
    $table->string('currency', 3);
    $table->string('status', 16)->index();        // pending | succeeded | failed | unknown
    $table->string('provider_payment_id')->nullable()->unique();
    $table->timestamps();
});
```

```php
// app/Services/Payments/ChargeOrder.php
<?php

declare(strict_types=1);

namespace App\Services\Payments;

use App\Models\Order;
use App\Models\Payment;
use Illuminate\Database\UniqueConstraintViolationException;
use Illuminate\Http\Client\ConnectionException;
use Illuminate\Support\Str;

final class ChargeOrder
{
    public function __construct(private PaymentGatewayClient $gateway) {}

    public function __invoke(Order $order, string $idempotencyKey): Payment
    {
        try {
            $payment = Payment::create([
                'idempotency_key' => $idempotencyKey,
                'order_id' => $order->id,
                'amount' => $order->total,
                'currency' => $order->currency,
                'status' => 'pending',
            ]);
        } catch (UniqueConstraintViolationException) {
            // Repeated click or job retry: the operation already exists
            return Payment::where('idempotency_key', $idempotencyKey)->firstOrFail();
        }

        try {
            $response = $this->gateway->charge($payment); // sends the Idempotency-Key
            $payment->update([
                'status' => $response['status'] === 'succeeded' ? 'succeeded' : 'failed',
                'provider_payment_id' => $response['id'],
            ]);
        } catch (ConnectionException) {
            // Outcome unknown: don't retry the charge, check the status later
            $payment->update(['status' => 'unknown']);
            ReconcilePayment::dispatch($payment->id)->delay(now()->addMinute());
        }

        return $payment;
    }
}
```

The idempotency key is created where the user's intent originates: for example, it is generated when the checkout page is opened and passed along in a hidden form field. That way a double click on the "Pay" button also collapses into a single operation.

The `ReconcilePayment` job asks the provider for the status by key or by `provider_payment_id` and moves the payment to `succeeded` or `failed`. Reconciliation isn't a hack; it's a mandatory part of any money integration: once a payment goes `unknown`, only the provider knows the truth.

---

<a id="webhooks"></a>
## Incoming webhooks and concurrent debits

Providers deliver webhooks at least once: the same event can arrive twice, hit two nodes simultaneously, or arrive before your synchronous response. The handler must be idempotent:

```php
// app/Http/Controllers/Webhooks/GatewayWebhookController.php
public function __invoke(GatewayWebhookRequest $request): Response
{
    $event = $request->validatedEvent(); // including signature verification

    DB::transaction(function () use ($event): void {
        // UNIQUE(provider, event_id): the second webhook inserts nothing
        $inserted = DB::table('processed_webhooks')->insertOrIgnore([
            'provider' => 'gateway',
            'event_id' => $event['id'],
            'created_at' => now(),
        ]);

        if ($inserted === 0) {
            return; // already processed
        }

        $payment = Payment::where('provider_payment_id', $event['payment_id'])
            ->lockForUpdate()
            ->firstOrFail();

        $payment->update(['status' => $event['status']]);
    });

    return response()->noContent();
}
```

An internal balance (a wallet, bonus points) also needs protection against a race between two simultaneous debits. `SELECT ... FOR UPDATE` serializes operations on a single row:

```php
// app/Services/Wallet/DebitWallet.php
DB::transaction(function () use ($walletId, $amount, $operationId): void {
    $wallet = Wallet::whereKey($walletId)->lockForUpdate()->firstOrFail();

    if ($wallet->balance < $amount) {
        throw new InsufficientFunds();
    }

    // UNIQUE(operation_id) in the ledger: one more barrier against retries
    $wallet->entries()->create(['operation_id' => $operationId, 'amount' => -$amount]);
    $wallet->decrement('balance', $amount);
});
```

For more on pessimistic and optimistic locking, see the article on [distributed transactions](database-and-distributed-transactions).

---

<a id="circuit-breaker"></a>
## Circuit Breaker: stop calling someone who doesn't answer

Retries and timeouts save an individual request. But if the provider is down for ten minutes, every request still burns its entire timeout budget and all its retries. After a series of failures, a circuit breaker **opens the circuit**: requests to the provider fail immediately or return a fallback, and every N seconds a single probe request is let through.

Laravel has no built-in circuit breaker for the HTTP client, but with atomic cache operations (Redis) you can write one in a few dozen lines:

```php
// app/Support/CircuitBreaker.php
<?php

declare(strict_types=1);

namespace App\Support;

use Closure;
use Illuminate\Support\Facades\Cache;
use RuntimeException;
use Throwable;

final class CircuitBreaker
{
    public function __construct(
        private string $name,
        private int $failureThreshold = 5,   // consecutive failures before opening
        private int $openSeconds = 30,       // pause before a probe request
    ) {}

    /**
     * @template T
     * @param Closure(): T $call
     * @return T
     */
    public function call(Closure $call): mixed
    {
        if (Cache::has($this->key('open'))) {
            // Let exactly one probe request through after the pause (half-open)
            if (! Cache::add($this->key('probe'), true, $this->openSeconds)) {
                throw new RuntimeException("Circuit {$this->name} is open");
            }
        }

        try {
            $result = $call();
        } catch (Throwable $e) {
            $failures = Cache::increment($this->key('failures'));
            if ($failures >= $this->failureThreshold) {
                Cache::put($this->key('open'), true, $this->openSeconds * 10);
            }
            throw $e;
        }

        Cache::forget($this->key('failures'));
        Cache::forget($this->key('open'));

        return $result;
    }

    private function key(string $suffix): string
    {
        return "circuit:{$this->name}:{$suffix}";
    }
}
```

```php
// Usage
$status = (new CircuitBreaker('gateway'))
    ->call(fn () => $gateway->paymentStatus($paymentId));
```

For queued jobs, Laravel already ships a close equivalent, the `ThrottlesExceptions` middleware: after a given number of exceptions, it delays subsequent attempts instead of hammering a service that's down:

```php
// app/Jobs/SyncGameRounds.php
use Illuminate\Queue\Middleware\ThrottlesExceptions;

public function middleware(): array
{
    // 10 exceptions → 5-minute pause (in Laravel 11+ the second argument is in seconds)
    return [(new ThrottlesExceptions(10, 5 * 60))->by('provider-games')];
}
```

> [!NOTE]
> **The state must be shared.** If the failure counter lives in process memory, each of 50 workers has its own circuit, and the provider receives 50 × threshold requests before all of them open. Store the state in Redis.

---

<a id="bulkheads"></a>
## Isolation: a separate queue for each provider

Even with timeouts, a slow provider can occupy every worker on a shared queue, and registration confirmation emails end up waiting behind a queue clogged with payment gateway calls. The **Bulkhead** pattern isolates resources:

```php
// app/Jobs/CapturePayment.php
public function __construct(public readonly int $paymentId)
{
    $this->onQueue('payments-gateway');
}
```

```ini
; /etc/supervisor/conf.d/queue-payments.conf
[program:queue-payments-gateway]
command=php /var/www/app/current/artisan queue:work redis --queue=payments-gateway --timeout=60
numprocs=4
```

Each provider gets its own queue and its own process limit. When one provider slows down, only its jobs pile up.

---

<a id="limitations"></a>
## Where this approach stops working

* **Not every provider supports idempotency keys.** Then your only protection is reconciliation: before retrying, ask the provider for the list of operations by your `order_id` or `reference`. If there's no such API either, retrying a money operation without a manual check is unacceptable.
* **Keys have a limited lifetime.** Stripe's, for example, last 24 hours. A retry two days later will be a new operation.
* **A circuit breaker can "protect" you from a healthy provider.** A threshold that's too low opens the circuit over a couple of random errors. Tune the threshold and window based on real error statistics, not guesswork.
* **Complexity grows.** `unknown` statuses, reconciliation jobs, a webhook log: this is all code that needs to be tested and monitored. For a weather API widget on the home page, a timeout and a cache are enough; everything else is for money and orders.

---

<a id="common-mistakes"></a>
## Common Mistakes

**1. Relying on the default timeout.**
Waiting 30 seconds on every request turns a provider outage into a full-site outage. Set `connectTimeout` and `timeout` explicitly.

**2. Retrying a charge after a timeout.**
A timeout means "unknown," not "failure." Without an idempotency key, a retry can charge the customer a second time.

**3. Retrying 4xx responses.**
An invalid request won't become valid on the second attempt. Retry only connection errors, `429`, and `5xx`.

**4. Retrying at three levels at once.**
The HTTP client, the job, and the calling service multiply the number of requests. Pick one level.

**5. Checking for duplicates with a `SELECT` instead of a UNIQUE constraint.**
Two concurrent requests will both see "no record" and both perform the operation. Only a unique index in the database protects you.

**6. An in-process circuit breaker.**
Each worker opens its own circuit independently. Keep the state in shared storage.

**7. No reconciliation.**
Payments stuck in `unknown` hang forever, and users write to support. Scheduled reconciliation is mandatory.

---

<a id="checklist"></a>
## Checklist

1. Every external call has explicit `connectTimeout` and `timeout` values, aligned with the number of workers.
2. Slow calls are moved out of the user's HTTP request into a queue.
3. Only idempotent operations are retried, and only on connection errors, `429`, and `5xx`, with exponential backoff and jitter.
4. Money operations get an idempotency key before the call; the key is stored under a UNIQUE constraint and passed to the provider.
5. A timeout moves the operation to `unknown`, which then triggers reconciliation.
6. Webhooks are processed idempotently: an event log with UNIQUE `(provider, event_id)`.
7. Debits from internal balances run under `lockForUpdate()`.
8. The circuit breaker keeps its state in Redis; queued jobs use `ThrottlesExceptions`.
9. Each provider has its own queue and process limit.

---

## Summary

An external API isn't a function; it's a distributed transaction with an unreliable participant. Timeouts, retries, and circuit breakers limit the damage from its failures, while idempotency and reconciliation make sure those failures don't turn into double charges. If, after a provider outage, you can answer "did the payment go through?" without logging into the provider's dashboard, your integration is designed correctly.

---

<a id="self-test-quiz"></a>
## Self-Test Quiz

### Question 1: A charge request failed with a timeout. What is the correct thing to do?
- A) Retry the charge immediately: a timeout means the money wasn't taken.
- B) Move the payment to the `unknown` status and find out the result by querying the provider for the status (or retry with the same idempotency key).
- C) Mark the payment as failed and ask the user to pay again.

<details>
<summary>Click to view the answer</summary>

**Answer: B**
A timeout doesn't tell you whether the operation was performed. A blind retry (A) or the user paying again (C) can lead to a double charge. The safe options are to check the status or to retry the request with the same idempotency key so the provider returns the result of the first attempt.
</details>

### Question 2: Why doesn't checking "does a payment with this key already exist?" with a `SELECT` before the `INSERT` protect against duplicates?
- A) `SELECT` is slower than a unique index.
- B) Two concurrent requests can both see that no record exists and both perform the operation. Only a UNIQUE constraint in the database protects reliably.
- C) Laravel caches `SELECT` results.

<details>
<summary>Click to view the answer</summary>

**Answer: B**
There's a race window between the check and the insert. The unique index is checked atomically by the database itself, so the second `INSERT` gets a `UniqueConstraintViolationException`, and the code returns the existing operation.
</details>

### Question 3: Why store circuit breaker state in Redis rather than in process memory?
- A) Redis is faster than a process's RAM.
- B) So that all workers and nodes see the same circuit state and stop sending requests at the same time, rather than each one individually after its own series of failures.
- C) Laravel can't store data in process memory.

<details>
<summary>Click to view the answer</summary>

**Answer: B**
With local state, each of dozens of workers has to hit the failure threshold on its own, and a provider that's down receives dozens of times more requests. Shared state opens the circuit for the whole system at once.
</details>