Laravel AI Agents: Streaming, Token Budgets &amp; Typed Output | Mohamed Said       [Skip to content](#main)  [ ![](https://cdn.msaied.com/01KT78WE565VEMM3PSNQAAB0MH.png) Mohamed SaidLaravel Backend Engineer ](https://www.msaied.com/public) - [Home](https://www.msaied.com/public)
- [Projects](https://www.msaied.com/public/projects)
- [Articles](https://www.msaied.com/public/articles)
- [Certificates](https://www.msaied.com/public/certificates)
- [About](https://www.msaied.com/public#about)

           [  Contact](https://www.msaied.com/public#contact) Menu 

Menu
----

Close 

 - [HomeStart here](https://www.msaied.com/public)
- [ProjectsCase studies](https://www.msaied.com/public/projects)
- [ArticlesEngineering notes](https://www.msaied.com/public/articles)
- [CertificatesCredentials](https://www.msaied.com/public/certificates)
- [AboutHow I work](https://www.msaied.com/public#about)
- [ContactGet in touch](https://www.msaied.com/public#contact)

  [Start a conversation](https://www.msaied.com/public#contact) [WhatsApp](https://wa.me/201094619204) [Email](mailto:hello@msaied.com) 

 1. [Home](https://www.msaied.com/public)
2. /
3. [Articles](https://www.msaied.com/public/articles)
4. /
5. Production AI Agents in Laravel: Streaming, Token Budgets, and Structured Output Contracts

 Production AI Agents in Laravel: Streaming, Token Budgets, and Structured Output Contracts
===========================================================================================

 Learn how to build reliable production AI agents in Laravel with streaming responses, enforced token budgets, and typed structured output contracts that survive real-world load.

 ![](https://cdn.msaied.com/01M22N44A70A5MC2S599JP0MPH.webp) [Mohamed Said](https://www.msaied.com/public#person) Published 7 Jul 2026 · Updated 7 Jul 2026 · 4 min read

ShareCopy linkCopied

 ![Production AI Agents in Laravel: Streaming, Token Budgets, and Structured Output Contracts](https://cdn.msaied.com/385/7cc838eef4b97e19a20cbbfd52799f52.png) 

  On this page +1. [The Gap Between Demo and Production](#the-gap-between-demo-and-production)
2. [1. Streaming Responses Over SSE](#1-streaming-responses-over-sse)
3. [2. Token Budget Middleware](#2-token-budget-middleware)
4. [3. Structured Output Contracts](#3-structured-output-contracts)
5. [Key Takeaways](#key-takeaways)

 The Gap Between Demo and Production
-----------------------------------

Most Laravel AI tutorials stop at `Http::post('https://api.openai.com/v1/chat/completions', [...])`. That works for a demo. It falls apart in production when a single runaway prompt burns your monthly token budget in an afternoon, a slow model response blocks a PHP-FPM worker for 45 seconds, or a hallucinated JSON blob crashes your downstream pipeline.

This article covers three concrete patterns that close that gap: **streaming with SSE**, **token budget middleware**, and **structured output contracts with readonly DTOs**.

---

1. Streaming Responses Over SSE
-------------------------------

OpenAI's streaming API sends `text/event-stream` chunks. Laravel's `StreamedResponse` lets you forward them to the browser without buffering the entire completion in memory.

```php
// routes/web.php
Route::get('/chat/stream', ChatStreamController::class);

// app/Http/Controllers/ChatStreamController.php
final class ChatStreamController
{
    public function __invoke(ChatRequest $request, AgentService $agent): StreamedResponse
    {
        return response()->stream(
            function () use ($request, $agent) {
                foreach ($agent->stream($request->validated('message')) as $chunk) {
                    echo "data: {$chunk}\n\n";
                    ob_flush();
                    flush();
                }
                echo "data: [DONE]\n\n";
            },
            200,
            [
                'Content-Type' => 'text/event-stream',
                'X-Accel-Buffering' => 'no', // critical for nginx
                'Cache-Control' => 'no-cache',
            ]
        );
    }
}

```

Inside `AgentService::stream()`, use a generator that reads the chunked HTTP response line by line:

```php
public function stream(string $prompt): \Generator
{
    $response = Http::withToken(config('services.openai.key'))
        ->withOptions(['stream' => true])
        ->post('https://api.openai.com/v1/chat/completions', [
            'model' => 'gpt-4o',
            'stream' => true,
            'messages' => [['role' => 'user', 'content' => $prompt]],
        ]);

    $body = $response->toPsrResponse()->getBody();

    while (! $body->eof()) {
        $line = trim($body->read(512));
        if (str_starts_with($line, 'data: ') && $line !== 'data: [DONE]') {
            $data = json_decode(substr($line, 6), true);
            $content = $data['choices'][0]['delta']['content'] ?? '';
            if ($content !== '') {
                yield $content;
            }
        }
    }
}

```

> **Note:** Set `fastcgi_buffering off` in nginx and ensure PHP-FPM output buffering is disabled for the relevant location block.

---

2. Token Budget Middleware
--------------------------

Token budgets belong at the pipeline layer, not scattered across controllers. A dedicated middleware class can inspect the request, estimate prompt tokens, and reject or downgrade the model before the HTTP call is made.

```php
final class EnforceTokenBudget
{
    // Rough heuristic: 1 token ≈ 4 chars
    private const CHARS_PER_TOKEN = 4;
    private const HARD_LIMIT = 8_000;

    public function handle(AgentPayload $payload, Closure $next): mixed
    {
        $estimated = (int) ceil(
            strlen($payload->systemPrompt . $payload->userMessage) / self::CHARS_PER_TOKEN
        );

        if ($estimated > self::HARD_LIMIT) {
            throw new TokenBudgetExceededException($estimated, self::HARD_LIMIT);
        }

        // Downgrade model for large-but-acceptable prompts
        if ($estimated > 4_000) {
            $payload = $payload->withModel('gpt-4o-mini');
        }

        return $next($payload);
    }
}

```

Wire it into a custom `AgentPipeline`:

```php
$result = app(Pipeline::class)
    ->send(new AgentPayload($system, $user))
    ->through([
        EnforceTokenBudget::class,
        SanitizeUserInput::class,
        InjectRagContext::class,
    ])
    ->thenReturn();

```

---

3. Structured Output Contracts
------------------------------

OpenAI's `response_format` with `json_schema` mode guarantees the model returns a specific shape. Pair it with a PHP 8.3 readonly DTO and a custom cast for end-to-end type safety.

```php
readonly final class ProductSummary
{
    public function __construct(
        public string $name,
        public string $oneLinePitch,
        /** @var list */
        public array $keyFeatures,
        public int $estimatedPriceUsd,
    ) {}

    public static function fromArray(array $data): self
    {
        return new self(
            name: $data['name'],
            oneLinePitch: $data['one_line_pitch'],
            keyFeatures: $data['key_features'],
            estimatedPriceUsd: $data['estimated_price_usd'],
        );
    }
}

```

Pass the schema to OpenAI:

```php
$response = Http::withToken(config('services.openai.key'))
    ->post('https://api.openai.com/v1/chat/completions', [
        'model' => 'gpt-4o',
        'response_format' => [
            'type' => 'json_schema',
            'json_schema' => [
                'name' => 'product_summary',
                'strict' => true,
                'schema' => [
                    'type' => 'object',
                    'properties' => [
                        'name' => ['type' => 'string'],
                        'one_line_pitch' => ['type' => 'string'],
                        'key_features' => ['type' => 'array', 'items' => ['type' => 'string']],
                        'estimated_price_usd' => ['type' => 'integer'],
                    ],
                    'required' => ['name', 'one_line_pitch', 'key_features', 'estimated_price_usd'],
                    'additionalProperties' => false,
                ],
            ],
        ],
        'messages' => [['role' => 'user', 'content' => $prompt]],
    ]);

$summary = ProductSummary::fromArray(
    json_decode($response->json('choices.0.message.content'), true)
);

```

Because `strict: true` is set, OpenAI will never return extra keys or omit required ones. Your `fromArray` factory becomes a safe assertion, not defensive guesswork.

---

Key Takeaways
-------------

- Use `response()->stream()` with a generator and `X-Accel-Buffering: no` for true SSE streaming in Laravel.
- Enforce token budgets in a pipeline middleware layer, not inside controllers or service methods.
- `json_schema` with `strict: true` eliminates hallucinated keys; map the response to a readonly DTO immediately.
- Estimate prompt tokens early (chars ÷ 4 is good enough for budget checks) to avoid surprise costs.
- Model downgrade logic belongs in the pipeline, keeping your agent service stateless and testable.

- [laravel](https://www.msaied.com/public/articles?search=laravel)
- [ai](https://www.msaied.com/public/articles?search=ai)
- [llm](https://www.msaied.com/public/articles?search=llm)
- [php](https://www.msaied.com/public/articles?search=php)

 Frequently asked questions 
---------------------------

  Does Laravel Octane break SSE streaming?Octane's Swoole driver handles streaming natively via its response object. Replace `response()-&gt;stream()` with `$server-&gt;push()` patterns or use the FrankenPHP driver, which supports streamed responses out of the box. Test with a real client; Octane workers do not buffer the way PHP-FPM does.

   How accurate is the chars-divided-by-4 token estimate?It is a conservative heuristic suitable for budget guards. For precise counts, use the tiktoken-php library or call OpenAI's tokenizer endpoint. The heuristic intentionally over-counts to provide a safety margin before hitting hard model context limits.

   Can I use structured output with streaming enabled simultaneously?Yes. OpenAI supports both `stream: true` and `response\_format: json\_schema` together. The JSON is streamed as partial tokens, so you must buffer the full stream before parsing the DTO. Only enable streaming for structured output if you need progress indication; otherwise a standard blocking call is simpler.

   ![Mohamed Said](https://cdn.msaied.com/01M22N44A70A5MC2S599JP0MPH.webp)About the author
----------------

[Mohamed Said](https://www.msaied.com/public#person)Senior Backend Engineer specializing in Laravel, scalable SaaS platforms, APIs, and cloud infrastructure. I build secure, high-performance web applications that help businesses grow.

[About](https://www.msaied.com/public#about) [GitHub ↗](https://github.com/EG-Mohamed) [LinkedIn ↗](https://www.linkedin.com/in/msaiedm/) [WhatsApp ↗](https://wa.me/201094619204) [Email Address ↗](mailto:hello@msaied.com) [My CV ↗](https://drive.google.com/file/u/0/d/1MF20IPRJyzfy32mhEutjL5EpSls0w2Q8/view)  

   [Previous articleFilament v3 to v4: Breaking Changes, Migration Patterns, and Refactor Strategies](https://www.msaied.com/public/articles/filament-v3-to-v4-breaking-changes-migration-patterns-and-refactor-strategies) [Next articleMulti-Tenant SaaS with Laravel + Filament: Scoped Queues, Notifications, and Storage per Tenant](https://www.msaied.com/public/articles/multi-tenant-saas-with-laravel-filament-scoped-queues-notifications-and-storage-per-tenant)  

   On this page
-------------

1. [The Gap Between Demo and Production](#the-gap-between-demo-and-production)
2. [1. Streaming Responses Over SSE](#1-streaming-responses-over-sse)
3. [2. Token Budget Middleware](#2-token-budget-middleware)
4. [3. Structured Output Contracts](#3-structured-output-contracts)
5. [Key Takeaways](#key-takeaways)

 ###  Have a technical challenge?

 Tell me what you’re building. I reply within two working days.

[Start a conversation](https://www.msaied.com/public#contact) 

   Related articles
-----------------

 [ ![](https://cdn.msaied.com/740/cce86edc21eddcbdd2f2454fadaf9c70.png)  · 3 min read### The Pipeline Pattern in Laravel: Custom Pipelines Beyond Middleware

5 Oct 2026 ](https://www.msaied.com/public/articles/the-pipeline-pattern-in-laravel-custom-pipelines-beyond-middleware-1) [ ![](https://cdn.msaied.com/739/2d6897fdcdcf090613f96f72a64b8a78.png)  · 4 min read### MySQL Full-Text Search in Laravel: Indexes, Relevance Scoring, and Boolean Mode

4 Oct 2026 ](https://www.msaied.com/public/articles/mysql-full-text-search-in-laravel-indexes-relevance-scoring-and-boolean-mode) [ ![](https://cdn.msaied.com/738/073696a3fefe18bec825beec5ac658f5.png)  · 4 min read### Laravel Queue Rate-Limited Middleware: Throttling Jobs Without Losing Work

4 Oct 2026 ](https://www.msaied.com/public/articles/laravel-queue-rate-limited-middleware-throttling-jobs-without-losing-work) 

  Have a technical challenge?
----------------------------

Tell me what you’re building. I reply within two working days.

 [Discuss your project ↗](https://www.msaied.com/public#contact) 

  © 2026 Mohamed Said · Built with Laravel, meant to last.Senior Backend Engineer specializing in Laravel, scalable SaaS platforms, APIs, and cloud infrastructure. I build secure, high-performance web applications that help businesses grow.

 - [Home](https://www.msaied.com/public)
- [Articles](https://www.msaied.com/public/articles)
- [Certificates](https://www.msaied.com/public/certificates)
- [GitHub](https://github.com/EG-Mohamed)
- [LinkedIn](https://www.linkedin.com/in/msaiedm/)
- [WhatsApp](https://wa.me/201094619204)
- [Email Address](mailto:hello@msaied.com)
- [My CV](https://drive.google.com/file/u/0/d/1MF20IPRJyzfy32mhEutjL5EpSls0w2Q8/view)
- [Sitemap](https://www.msaied.com/public/sitemap.xml)
