Laravel AI Agents: Streaming &amp; Structured Output | Mohamed Said       [Skip to content](#main)  [ ![](https://cdn.msaied.com/01KT78WE565VEMM3PSNQAAB0MH.png) Mohamed SaidLaravel Backend Engineer ](https://www.msaied.com/public) - [Home](https://www.msaied.com/public)
- [Projects](https://www.msaied.com/public/projects)
- [Articles](https://www.msaied.com/public/articles)
- [Certificates](https://www.msaied.com/public/certificates)
- [About](https://www.msaied.com/public#about)

           [  Contact](https://www.msaied.com/public#contact) Menu 

Menu
----

Close 

 - [HomeStart here](https://www.msaied.com/public)
- [ProjectsCase studies](https://www.msaied.com/public/projects)
- [ArticlesEngineering notes](https://www.msaied.com/public/articles)
- [CertificatesCredentials](https://www.msaied.com/public/certificates)
- [AboutHow I work](https://www.msaied.com/public#about)
- [ContactGet in touch](https://www.msaied.com/public#contact)

  [Start a conversation](https://www.msaied.com/public#contact) [WhatsApp](https://wa.me/201094619204) [Email](mailto:hello@msaied.com) 

 1. [Home](https://www.msaied.com/public)
2. /
3. [Articles](https://www.msaied.com/public/articles)
4. /
5. Production AI Agents in Laravel: Streaming, Token Budgets, and Structured Output Contracts

 Production AI Agents in Laravel: Streaming, Token Budgets, and Structured Output Contracts
===========================================================================================

 Building reliable AI agents in Laravel means more than wiring up an API call. Learn how to stream responses safely, enforce token budgets, and lock down structured output with typed contracts your application can trust.

 ![](https://cdn.msaied.com/01M22N44A70A5MC2S599JP0MPH.webp) [Mohamed Said](https://www.msaied.com/public#person) Published 26 Sep 2026 · Updated 26 Sep 2026 · 3 min read

ShareCopy linkCopied

 ![Production AI Agents in Laravel: Streaming, Token Budgets, and Structured Output Contracts](https://cdn.msaied.com/704/165bcb76b898fe9911130d2c93b9b805.png) 

  On this page +1. [The Gap Between a Demo and a Production Agent](#the-gap-between-a-demo-and-a-production-agent)
2. [1. Server-Sent Events Streaming with Laravel](#1-server-sent-events-streaming-with-laravel)
3. [2. Enforcing Token Budgets Before the Request Leaves](#2-enforcing-token-budgets-before-the-request-leaves)
4. [3. Typed Structured Output Contracts](#3-typed-structured-output-contracts)
5. [Takeaways](#takeaways)

 The Gap Between a Demo and a Production Agent
---------------------------------------------

Most Laravel AI tutorials stop at `Http::post('https://api.openai.com/v1/chat/completions', [...])` and call it done. In production you face three hard problems: responses that take 30+ seconds must stream to the browser, runaway prompts blow your cost budget, and untyped JSON blobs from the model break your downstream logic silently. This article solves all three.

---

1. Server-Sent Events Streaming with Laravel
--------------------------------------------

OpenAI's streaming endpoint sends newline-delimited `data:` chunks. Laravel's `StreamedResponse` is the right primitive.

```php
// routes/web.php
Route::get('/agent/stream', AgentStreamController::class);

```

```php
final class AgentStreamController
{
    public function __invoke(Request $request): StreamedResponse
    {
        $messages = $request->validate(['messages' => 'required|array']);

        return response()->stream(function () use ($messages) {
            $stream = OpenAI::chat()->createStreamed([
                'model'    => 'gpt-4o',
                'messages' => $messages['messages'],
            ]);

            foreach ($stream as $response) {
                $delta = $response->choices[0]->delta->content ?? '';
                if ($delta !== '') {
                    echo 'data: ' . json_encode(['token' => $delta]) . "\n\n";
                    ob_flush();
                    flush();
                }
            }

            echo "data: [DONE]\n\n";
            ob_flush();
            flush();
        }, 200, [
            'Content-Type'      => 'text/event-stream',
            'Cache-Control'     => 'no-cache',
            'X-Accel-Buffering' => 'no', // critical for Nginx
        ]);
    }
}

```

The `X-Accel-Buffering: no` header is the most commonly forgotten detail. Without it, Nginx buffers the entire response before forwarding it.

---

2. Enforcing Token Budgets Before the Request Leaves
----------------------------------------------------

Token overruns are a billing and latency problem. Enforce a budget at the call site, not as an afterthought.

```php
final class TokenBudget
{
    public function __construct(
        private readonly int $maxPromptTokens = 3_000,
        private readonly int $maxCompletionTokens = 1_000,
    ) {}

    /** Rough estimate: 1 token ≈ 4 chars for English prose */
    public function assertPromptFits(array $messages): void
    {
        $chars = array_sum(array_map(
            fn($m) => strlen($m['content'] ?? ''),
            $messages
        ));

        $estimated = (int) ceil($chars / 4);

        if ($estimated > $this->maxPromptTokens) {
            throw new PromptTooLargeException(
                "Estimated {$estimated} tokens exceeds budget of {$this->maxPromptTokens}."
            );
        }
    }

    public function completionLimit(): int
    {
        return $this->maxCompletionTokens;
    }
}

```

Bind it as a singleton and inject it into your agent service. Pass `max_tokens` explicitly on every API call — never leave it open-ended in production.

```php
$budget->assertPromptFits($messages);

$response = OpenAI::chat()->create([
    'model'      => 'gpt-4o',
    'messages'   => $messages,
    'max_tokens' => $budget->completionLimit(),
]);

```

---

3. Typed Structured Output Contracts
------------------------------------

OpenAI's JSON mode and structured outputs return a string you must decode. Wrap that decode in a typed DTO so a schema mismatch throws immediately rather than propagating a null through your domain.

```php
final readonly class ExtractedLeadData
{
    public function __construct(
        public string $companyName,
        public string $contactEmail,
        public ?string $phoneNumber,
        public int $estimatedEmployees,
    ) {}

    public static function fromModelResponse(string $json): self
    {
        $data = json_decode($json, true, flags: JSON_THROW_ON_ERROR);

        return new self(
            companyName:         $data['company_name'] ?? throw new MalformedAgentResponseException('company_name'),
            contactEmail:        filter_var($data['contact_email'] ?? '', FILTER_VALIDATE_EMAIL)
                                     ?: throw new MalformedAgentResponseException('contact_email'),
            phoneNumber:         $data['phone_number'] ?? null,
            estimatedEmployees:  (int) ($data['estimated_employees']
                                     ?? throw new MalformedAgentResponseException('estimated_employees')),
        );
    }
}

```

Pair this with a `response_format: { type: 'json_object' }` parameter and a system prompt that describes the exact schema. The DTO constructor acts as your runtime contract — if the model drifts, you catch it at the boundary.

---

Takeaways
---------

- Set `X-Accel-Buffering: no` on every SSE response or Nginx will silently buffer your stream.
- Estimate token counts before the request and pass `max_tokens` explicitly — never leave it unbounded.
- Decode model JSON into typed DTOs at the boundary; let the constructor throw on schema violations.
- Treat `MalformedAgentResponseException` as a retryable error — models occasionally produce malformed JSON even in JSON mode.
- Keep `TokenBudget` as a named singleton so budget policy is one config change, not a grep-and-replace.

- [laravel](https://www.msaied.com/public/articles?search=laravel)
- [ai](https://www.msaied.com/public/articles?search=ai)
- [llm](https://www.msaied.com/public/articles?search=llm)
- [php](https://www.msaied.com/public/articles?search=php)

 Frequently asked questions 
---------------------------

  Why does my streamed response appear all at once in the browser despite using StreamedResponse?Almost always it is Nginx buffering the response. Add the `X-Accel-Buffering: no` response header and ensure `fastcgi\_buffering off` is not overriding it at the server block level. Also confirm `ob\_flush()` and `flush()` are called after each chunk.

   Is the 1 token ≈ 4 characters estimate reliable enough for a budget guard?It is a conservative heuristic for English text. For multilingual content or code, tokens are denser and the estimate will under-count. Use it as a pre-flight safety check, not a billing-accurate counter. For precise counts, integrate a tiktoken-compatible tokenizer library.

   Should I retry when ExtractedLeadData::fromModelResponse throws MalformedAgentResponseException?Yes, with a low retry limit (1–2 attempts) and a stricter system prompt on the retry. Log the raw JSON on failure so you can audit model drift over time. If failures are frequent, switch to OpenAI's structured outputs feature which enforces your JSON schema server-side.

   ![Mohamed Said](https://cdn.msaied.com/01M22N44A70A5MC2S599JP0MPH.webp)About the author
----------------

[Mohamed Said](https://www.msaied.com/public#person)Senior Backend Engineer specializing in Laravel, scalable SaaS platforms, APIs, and cloud infrastructure. I build secure, high-performance web applications that help businesses grow.

[About](https://www.msaied.com/public#about) [GitHub ↗](https://github.com/EG-Mohamed) [LinkedIn ↗](https://www.linkedin.com/in/msaiedm/) [WhatsApp ↗](https://wa.me/201094619204) [Email Address ↗](mailto:hello@msaied.com) [My CV ↗](https://drive.google.com/file/u/0/d/1MF20IPRJyzfy32mhEutjL5EpSls0w2Q8/view)  

   [Previous articleDecide with Jev: Build a Laravel AI Content Preflight Checker That Returns a Probability](https://www.msaied.com/public/articles/decide-with-jev-build-a-laravel-ai-content-preflight-checker-that-returns-a-probability) [Next articleOctane State Leakage: Detecting and Fixing Shared-Memory Bugs in Laravel Workers](https://www.msaied.com/public/articles/octane-state-leakage-detecting-and-fixing-shared-memory-bugs-in-laravel-workers)  

   On this page
-------------

1. [The Gap Between a Demo and a Production Agent](#the-gap-between-a-demo-and-a-production-agent)
2. [1. Server-Sent Events Streaming with Laravel](#1-server-sent-events-streaming-with-laravel)
3. [2. Enforcing Token Budgets Before the Request Leaves](#2-enforcing-token-budgets-before-the-request-leaves)
4. [3. Typed Structured Output Contracts](#3-typed-structured-output-contracts)
5. [Takeaways](#takeaways)

 ###  Have a technical challenge?

 Tell me what you’re building. I reply within two working days.

[Start a conversation](https://www.msaied.com/public#contact) 

   Related articles
-----------------

 [ ![](https://cdn.msaied.com/740/cce86edc21eddcbdd2f2454fadaf9c70.png)  · 3 min read### The Pipeline Pattern in Laravel: Custom Pipelines Beyond Middleware

5 Oct 2026 ](https://www.msaied.com/public/articles/the-pipeline-pattern-in-laravel-custom-pipelines-beyond-middleware-1) [ ![](https://cdn.msaied.com/739/2d6897fdcdcf090613f96f72a64b8a78.png)  · 4 min read### MySQL Full-Text Search in Laravel: Indexes, Relevance Scoring, and Boolean Mode

4 Oct 2026 ](https://www.msaied.com/public/articles/mysql-full-text-search-in-laravel-indexes-relevance-scoring-and-boolean-mode) [ ![](https://cdn.msaied.com/738/073696a3fefe18bec825beec5ac658f5.png)  · 4 min read### Laravel Queue Rate-Limited Middleware: Throttling Jobs Without Losing Work

4 Oct 2026 ](https://www.msaied.com/public/articles/laravel-queue-rate-limited-middleware-throttling-jobs-without-losing-work) 

  Have a technical challenge?
----------------------------

Tell me what you’re building. I reply within two working days.

 [Discuss your project ↗](https://www.msaied.com/public#contact) 

  © 2026 Mohamed Said · Built with Laravel, meant to last.Senior Backend Engineer specializing in Laravel, scalable SaaS platforms, APIs, and cloud infrastructure. I build secure, high-performance web applications that help businesses grow.

 - [Home](https://www.msaied.com/public)
- [Articles](https://www.msaied.com/public/articles)
- [Certificates](https://www.msaied.com/public/certificates)
- [GitHub](https://github.com/EG-Mohamed)
- [LinkedIn](https://www.linkedin.com/in/msaiedm/)
- [WhatsApp](https://wa.me/201094619204)
- [Email Address](mailto:hello@msaied.com)
- [My CV](https://drive.google.com/file/u/0/d/1MF20IPRJyzfy32mhEutjL5EpSls0w2Q8/view)
- [Sitemap](https://www.msaied.com/public/sitemap.xml)
