Laravel AI Agents: Streaming &amp; Structured Output | Mohamed Said       [Skip to content](#main)  [ ![](https://cdn.msaied.com/01KT78WE565VEMM3PSNQAAB0MH.png) Mohamed SaidLaravel Backend Engineer ](https://www.msaied.com/public) - [Home](https://www.msaied.com/public)
- [Projects](https://www.msaied.com/public/projects)
- [Articles](https://www.msaied.com/public/articles)
- [Certificates](https://www.msaied.com/public/certificates)
- [About](https://www.msaied.com/public#about)

           [  Contact](https://www.msaied.com/public#contact) Menu 

Menu
----

Close 

 - [HomeStart here](https://www.msaied.com/public)
- [ProjectsCase studies](https://www.msaied.com/public/projects)
- [ArticlesEngineering notes](https://www.msaied.com/public/articles)
- [CertificatesCredentials](https://www.msaied.com/public/certificates)
- [AboutHow I work](https://www.msaied.com/public#about)
- [ContactGet in touch](https://www.msaied.com/public#contact)

  [Start a conversation](https://www.msaied.com/public#contact) [WhatsApp](https://wa.me/201094619204) [Email](mailto:hello@msaied.com) 

 1. [Home](https://www.msaied.com/public)
2. /
3. [Articles](https://www.msaied.com/public/articles)
4. /
5. Production AI Agents in Laravel: Streaming, Token Budgets, and Structured Output Contracts

 Production AI Agents in Laravel: Streaming, Token Budgets, and Structured Output Contracts
===========================================================================================

 Building reliable AI agents means more than wiring up an API key. Learn how to stream responses safely, enforce token budgets, and lock down structured output contracts so your Laravel agents behave predictably in production.

 ![](https://cdn.msaied.com/01M22N44A70A5MC2S599JP0MPH.webp) [Mohamed Said](https://www.msaied.com/public#person) Published 30 Jun 2026 · Updated 30 Jun 2026 · 4 min read

ShareCopy linkCopied

 ![Production AI Agents in Laravel: Streaming, Token Budgets, and Structured Output Contracts](https://cdn.msaied.com/325/ee4e73c7fdfa922304594beaac92c93c.png) 

  On this page +1. [Production AI Agents in Laravel: Streaming, Token Budgets, and Structured Output Contracts](#production-ai-agents-in-laravel-streaming-token-budgets-and-structured-output-contracts)
2. [Streaming Responses Over SSE](#streaming-responses-over-sse)
3. [Enforcing Token Budgets](#enforcing-token-budgets)
4. [Structured Output Contracts](#structured-output-contracts)
5. [Putting It Together in a Job](#putting-it-together-in-a-job)
6. [Takeaways](#takeaways)

 Production AI Agents in Laravel: Streaming, Token Budgets, and Structured Output Contracts
------------------------------------------------------------------------------------------

Most tutorials stop at "call the API, dump the response." Production agents are a different beast. You need to stream tokens to the browser without blocking a PHP worker for 30 seconds, enforce hard token budgets so a runaway prompt doesn't drain your quota, and guarantee that the model returns data your application can actually parse. This article covers all three.

---

### Streaming Responses Over SSE

Laravel's `StreamedResponse` pairs naturally with OpenAI's streaming API. The key is flushing output incrementally without buffering the entire completion.

```php
use Illuminate\Support\Facades\Route;
use Symfony\Component\HttpFoundation\StreamedResponse;
use OpenAI\Laravel\Facades\OpenAI;

Route::get('/chat/stream', function () {
    return new StreamedResponse(function () {
        $stream = OpenAI::chat()->createStreamed([
            'model' => 'gpt-4o',
            'messages' => [['role' => 'user', 'content' => request('prompt')]],
        ]);

        foreach ($stream as $response) {
            $delta = $response->choices[0]->delta->content ?? '';
            if ($delta !== '') {
                echo 'data: ' . json_encode(['token' => $delta]) . "\n\n";
                ob_flush();
                flush();
            }
        }

        echo "data: [DONE]\n\n";
    }, 200, [
        'Content-Type' => 'text/event-stream',
        'Cache-Control' => 'no-cache',
        'X-Accel-Buffering' => 'no', // critical for nginx
    ]);
});

```

`X-Accel-Buffering: no` is the header most people forget. Without it, nginx will buffer the entire response before forwarding it to the client, defeating the purpose of streaming entirely.

---

### Enforcing Token Budgets

Token overruns are a billing and latency problem. Enforce budgets at two layers: before the request (prompt token estimation) and inside the request (`max_tokens`).

```php
final class TokenBudget
{
    public function __construct(
        private readonly int $maxPromptTokens = 3_000,
        private readonly int $maxCompletionTokens = 1_000,
    ) {}

    public function assertPromptFits(string $prompt): void
    {
        // ~4 chars per token is a safe heuristic for English text
        $estimated = (int) ceil(mb_strlen($prompt) / 4);

        if ($estimated > $this->maxPromptTokens) {
            throw new \OverflowException(
                "Prompt exceeds budget: ~{$estimated} tokens (max {$this->maxPromptTokens})"
            );
        }
    }

    public function completionLimit(): int
    {
        return $this->maxCompletionTokens;
    }
}

```

Bind this as a singleton scoped to the current tenant or user plan:

```php
$this->app->scoped(TokenBudget::class, function () {
    $plan = auth()->user()?->plan ?? 'free';
    return match ($plan) {
        'pro'  => new TokenBudget(8_000, 2_000),
        default => new TokenBudget(3_000, 500),
    };
});

```

Using `scoped` rather than `singleton` ensures the budget resets per request, which matters under Octane.

---

### Structured Output Contracts

Asking a model to "return JSON" is not a contract. OpenAI's `response_format` with `json_schema` mode (available on `gpt-4o` and later) lets you enforce a schema server-side. Pair it with a DTO and a Pest assertion.

```php
$response = OpenAI::chat()->create([
    'model' => 'gpt-4o',
    'messages' => [
        ['role' => 'system', 'content' => 'Extract the invoice fields.'],
        ['role' => 'user', 'content' => $rawText],
    ],
    'response_format' => [
        'type' => 'json_schema',
        'json_schema' => [
            'name' => 'invoice',
            'strict' => true,
            'schema' => [
                'type' => 'object',
                'properties' => [
                    'vendor'  => ['type' => 'string'],
                    'amount'  => ['type' => 'number'],
                    'due_date'=> ['type' => 'string', 'format' => 'date'],
                ],
                'required' => ['vendor', 'amount', 'due_date'],
                'additionalProperties' => false,
            ],
        ],
    ],
]);

$data = json_decode($response->choices[0]->message->content, true, flags: JSON_THROW_ON_ERROR);
$invoice = InvoiceData::from($data); // Spatie Data DTO

```

With `strict: true`, the model will refuse to produce output that violates the schema rather than hallucinating extra fields. Validate the DTO immediately after hydration — never trust the model's output downstream without a type check.

---

### Putting It Together in a Job

For non-interactive agents, run the completion inside a queued job with a timeout that matches your token budget:

```php
class ExtractInvoiceJob implements ShouldQueue
{
    public int $timeout = 60;
    public int $tries = 2;

    public function handle(TokenBudget $budget, InvoiceExtractor $extractor): void
    {
        $budget->assertPromptFits($this->rawText);
        $invoice = $extractor->extract($this->rawText, $budget->completionLimit());
        InvoiceExtracted::dispatch($invoice);
    }
}

```

Set `$timeout` conservatively. A 500-token completion at peak load can still take 20+ seconds.

---

### Takeaways

- Add `X-Accel-Buffering: no` to every SSE response or nginx will swallow your stream.
- Use `scoped()` for per-request token budgets under Octane, not `singleton()`.
- OpenAI's `json_schema` response format with `strict: true` is a real contract, not a prompt suggestion.
- Validate and hydrate into a typed DTO immediately — never pass raw model output into business logic.
- Set explicit job `$timeout` values that reflect your worst-case token budget, not a generic default.

- [laravel](https://www.msaied.com/public/articles?search=laravel)
- [ai](https://www.msaied.com/public/articles?search=ai)
- [openai](https://www.msaied.com/public/articles?search=openai)
- [production](https://www.msaied.com/public/articles?search=production)

 Frequently asked questions 
---------------------------

  Why does my nginx proxy buffer the SSE stream even with StreamedResponse?Nginx buffers proxy responses by default. Set the `X-Accel-Buffering: no` response header to instruct nginx to pass chunks through immediately. You may also need `proxy\_buffering off` in your nginx config for non-Accel setups.

   Is the 4-characters-per-token heuristic accurate enough for budget enforcement?It is a safe overestimate for English prose, which is intentional. For precise counts use a tokenizer library (e.g., tiktoken via a PHP FFI binding), but the heuristic is sufficient for a pre-flight guard that errs on the side of caution.

   Does `json\_schema` response format work with all OpenAI models?Structured output with `strict: true` requires `gpt-4o` (2024-08-06 snapshot or later) or `gpt-4o-mini`. Earlier models support `response\_format: {type: json\_object}` but without schema enforcement, so the model can still produce non-conforming output.

   ![Mohamed Said](https://cdn.msaied.com/01M22N44A70A5MC2S599JP0MPH.webp)About the author
----------------

[Mohamed Said](https://www.msaied.com/public#person)Senior Backend Engineer specializing in Laravel, scalable SaaS platforms, APIs, and cloud infrastructure. I build secure, high-performance web applications that help businesses grow.

[About](https://www.msaied.com/public#about) [GitHub ↗](https://github.com/EG-Mohamed) [LinkedIn ↗](https://www.linkedin.com/in/msaiedm/) [WhatsApp ↗](https://wa.me/201094619204) [Email Address ↗](mailto:hello@msaied.com) [My CV ↗](https://drive.google.com/file/u/0/d/1MF20IPRJyzfy32mhEutjL5EpSls0w2Q8/view)  

   [Previous articleFilament v3 to v4 Migration: Breaking Changes and Practical Refactor Patterns](https://www.msaied.com/public/articles/filament-v3-to-v4-migration-breaking-changes-and-practical-refactor-patterns-1) [Next articleMulti-Tenant SaaS in Laravel: Isolating Tenant State with Scoped Service Bindings](https://www.msaied.com/public/articles/multi-tenant-saas-in-laravel-isolating-tenant-state-with-scoped-service-bindings)  

   On this page
-------------

1. [Production AI Agents in Laravel: Streaming, Token Budgets, and Structured Output Contracts](#production-ai-agents-in-laravel-streaming-token-budgets-and-structured-output-contracts)
2. [Streaming Responses Over SSE](#streaming-responses-over-sse)
3. [Enforcing Token Budgets](#enforcing-token-budgets)
4. [Structured Output Contracts](#structured-output-contracts)
5. [Putting It Together in a Job](#putting-it-together-in-a-job)
6. [Takeaways](#takeaways)

 ###  Have a technical challenge?

 Tell me what you’re building. I reply within two working days.

[Start a conversation](https://www.msaied.com/public#contact) 

   Related articles
-----------------

 [ ![](https://cdn.msaied.com/740/cce86edc21eddcbdd2f2454fadaf9c70.png)  · 3 min read### The Pipeline Pattern in Laravel: Custom Pipelines Beyond Middleware

5 Oct 2026 ](https://www.msaied.com/public/articles/the-pipeline-pattern-in-laravel-custom-pipelines-beyond-middleware-1) [ ![](https://cdn.msaied.com/739/2d6897fdcdcf090613f96f72a64b8a78.png)  · 4 min read### MySQL Full-Text Search in Laravel: Indexes, Relevance Scoring, and Boolean Mode

4 Oct 2026 ](https://www.msaied.com/public/articles/mysql-full-text-search-in-laravel-indexes-relevance-scoring-and-boolean-mode) [ ![](https://cdn.msaied.com/738/073696a3fefe18bec825beec5ac658f5.png)  · 4 min read### Laravel Queue Rate-Limited Middleware: Throttling Jobs Without Losing Work

4 Oct 2026 ](https://www.msaied.com/public/articles/laravel-queue-rate-limited-middleware-throttling-jobs-without-losing-work) 

  Have a technical challenge?
----------------------------

Tell me what you’re building. I reply within two working days.

 [Discuss your project ↗](https://www.msaied.com/public#contact) 

  © 2026 Mohamed Said · Built with Laravel, meant to last.Senior Backend Engineer specializing in Laravel, scalable SaaS platforms, APIs, and cloud infrastructure. I build secure, high-performance web applications that help businesses grow.

 - [Home](https://www.msaied.com/public)
- [Articles](https://www.msaied.com/public/articles)
- [Certificates](https://www.msaied.com/public/certificates)
- [GitHub](https://github.com/EG-Mohamed)
- [LinkedIn](https://www.linkedin.com/in/msaiedm/)
- [WhatsApp](https://wa.me/201094619204)
- [Email Address](mailto:hello@msaied.com)
- [My CV](https://drive.google.com/file/u/0/d/1MF20IPRJyzfy32mhEutjL5EpSls0w2Q8/view)
- [Sitemap](https://www.msaied.com/public/sitemap.xml)
