Laravel AI: Streaming, Token Budgets &amp; Structured Output | Mohamed Said       [Skip to content](#main)  [ ![](https://cdn.msaied.com/01KT78WE565VEMM3PSNQAAB0MH.png) Mohamed SaidLaravel Backend Engineer ](https://www.msaied.com/public) - [Home](https://www.msaied.com/public)
- [Projects](https://www.msaied.com/public/projects)
- [Articles](https://www.msaied.com/public/articles)
- [Certificates](https://www.msaied.com/public/certificates)
- [About](https://www.msaied.com/public#about)

           [  Contact](https://www.msaied.com/public#contact) Menu 

Menu
----

Close 

 - [HomeStart here](https://www.msaied.com/public)
- [ProjectsCase studies](https://www.msaied.com/public/projects)
- [ArticlesEngineering notes](https://www.msaied.com/public/articles)
- [CertificatesCredentials](https://www.msaied.com/public/certificates)
- [AboutHow I work](https://www.msaied.com/public#about)
- [ContactGet in touch](https://www.msaied.com/public#contact)

  [Start a conversation](https://www.msaied.com/public#contact) [WhatsApp](https://wa.me/201094619204) [Email](mailto:hello@msaied.com) 

 1. [Home](https://www.msaied.com/public)
2. /
3. [Articles](https://www.msaied.com/public/articles)
4. /
5. Streaming AI Responses in Laravel: Token Budgets, Structured Output, and Agent Contracts

 Streaming AI Responses in Laravel: Token Budgets, Structured Output, and Agent Contracts
=========================================================================================

 Learn how to stream LLM responses in Laravel, enforce token budgets, and lock structured output to typed PHP contracts — keeping production AI agents predictable and cost-controlled.

 ![](https://cdn.msaied.com/01M22N44A70A5MC2S599JP0MPH.webp) [Mohamed Said](https://www.msaied.com/public#person) Published 23 Aug 2026 · Updated 23 Aug 2026 · 4 min read

ShareCopy linkCopied

 ![Streaming AI Responses in Laravel: Token Budgets, Structured Output, and Agent Contracts](https://cdn.msaied.com/584/3ffe135721c65ab3f9b40401dc3c41de.png) 

  On this page +1. [The Problem With Naive LLM Integration](#the-problem-with-naive-llm-integration)
2. [Streaming Responses to the Browser](#streaming-responses-to-the-browser)
3. [Enforcing Token Budgets](#enforcing-token-budgets)
4. [1. Hard limit via max\_tokens](#1-hard-limit-via-codemax-tokenscode)
5. [2. Prompt token pre-check](#2-prompt-token-pre-check)
6. [Structured Output Contracts](#structured-output-contracts)
7. [Wiring It Together in a Job](#wiring-it-together-in-a-job)
8. [Key Takeaways](#key-takeaways)

 The Problem With Naive LLM Integration
--------------------------------------

Most Laravel + LLM tutorials show a single `chat()` call and a `dd($response->content)`. That works in a demo. In production you face three hard problems: responses block the HTTP worker until the model finishes, uncapped prompts silently drain your budget, and free-form JSON from the model breaks your downstream code without warning.

This article tackles all three with concrete patterns.

---

Streaming Responses to the Browser
----------------------------------

OpenAI's streaming API sends server-sent events (SSE). Laravel's `StreamedResponse` lets you forward them without buffering the entire completion.

```php
use Illuminate\Http\Response;
use OpenAI\Laravel\Facades\OpenAI;

Route::get('/chat', function () {
    return response()->stream(function () {
        $stream = OpenAI::chat()->createStreamed([
            'model' => 'gpt-4o',
            'messages' => [['role' => 'user', 'content' => request('prompt')]],
        ]);

        foreach ($stream as $response) {
            $delta = $response->choices[0]->delta->content ?? '';
            if ($delta !== '') {
                echo "data: " . json_encode(['token' => $delta]) . "\n\n";
                ob_flush();
                flush();
            }
        }

        echo "data: [DONE]\n\n";
    }, 200, [
        'Content-Type' => 'text/event-stream',
        'X-Accel-Buffering' => 'no', // critical for nginx
        'Cache-Control' => 'no-cache',
    ]);
});

```

The `X-Accel-Buffering: no` header is the most commonly forgotten detail when running behind nginx — without it, nginx buffers the entire stream before forwarding.

---

Enforcing Token Budgets
-----------------------

Token costs compound fast when users craft adversarial prompts. Enforce budgets at two layers.

### 1. Hard limit via `max_tokens`

```php
$payload = [
    'model' => 'gpt-4o',
    'max_tokens' => config('ai.max_completion_tokens', 512),
    'messages' => $messages,
];

```

### 2. Prompt token pre-check

Count tokens before sending using `tiktoken-php` or a simple heuristic, and reject early:

```php
use Yethee\Tiktoken\EncoderProvider;

final class TokenBudgetGuard
{
    private const MODEL_LIMIT = 8_192;
    private const RESERVED_FOR_COMPLETION = 512;

    public function __construct(private EncoderProvider $provider) {}

    public function assertFits(array $messages, string $model = 'gpt-4o'): void
    {
        $encoder = $this->provider->getForModel($model);
        $tokens = array_sum(
            array_map(fn ($m) => count($encoder->encode($m['content'])), $messages)
        );

        $budget = self::MODEL_LIMIT - self::RESERVED_FOR_COMPLETION;

        if ($tokens > $budget) {
            throw new TokenBudgetExceededException($tokens, $budget);
        }
    }
}

```

Bind this as a singleton and inject it into your agent service. Throw early — never let an oversized prompt reach the API.

---

Structured Output Contracts
---------------------------

OpenAI's `response_format` with `json_schema` mode guarantees the model returns JSON matching your schema. Pair that with a typed PHP DTO and you get end-to-end type safety.

```php
readonly class ProductSuggestion
{
    public function __construct(
        public string $name,
        public string $reason,
        public int $confidencePercent,
    ) {}

    public static function fromArray(array $data): self
    {
        return new self(
            name: $data['name'],
            reason: $data['reason'],
            confidencePercent: $data['confidence_percent'],
        );
    }
}

```

```php
$response = OpenAI::chat()->create([
    'model' => 'gpt-4o-2024-08-06', // structured output requires this or later
    'messages' => $messages,
    'response_format' => [
        'type' => 'json_schema',
        'json_schema' => [
            'name' => 'product_suggestion',
            'strict' => true,
            'schema' => [
                'type' => 'object',
                'properties' => [
                    'name' => ['type' => 'string'],
                    'reason' => ['type' => 'string'],
                    'confidence_percent' => ['type' => 'integer'],
                ],
                'required' => ['name', 'reason', 'confidence_percent'],
                'additionalProperties' => false,
            ],
        ],
    ],
]);

$suggestion = ProductSuggestion::fromArray(
    json_decode($response->choices[0]->message->content, true, flags: JSON_THROW_ON_ERROR)
);

```

With `strict: true` the model will refuse to emit keys not in your schema. Validation failures become model refusals, not silent bad data.

---

Wiring It Together in a Job
---------------------------

For non-interactive workloads, push the agent call to a queued job and store the result:

```php
class RunProductSuggestionAgent implements ShouldQueue
{
    use Dispatchable, Queueable;

    public int $tries = 2;
    public int $timeout = 60;

    public function __construct(private int $productId) {}

    public function handle(TokenBudgetGuard $guard, ProductRepository $repo): void
    {
        $product = $repo->findOrFail($this->productId);
        $messages = MessageBuilder::forProduct($product);

        $guard->assertFits($messages);

        // ... call OpenAI, hydrate DTO, persist
    }
}

```

Set `$timeout` explicitly — the default 60 s is often too short for large completions and too long to leave zombie workers hanging.

---

Key Takeaways
-------------

- Stream via `response()->stream()` and set `X-Accel-Buffering: no` for nginx.
- Pre-check prompt token counts before hitting the API; throw early.
- Use `max_tokens` as a hard ceiling on every request.
- Lock structured output with `json_schema` + `strict: true` and hydrate into readonly DTOs.
- Push long-running completions to queued jobs with explicit `$timeout` and `$tries`.

- [laravel](https://www.msaied.com/public/articles?search=laravel)
- [ai](https://www.msaied.com/public/articles?search=ai)
- [llm](https://www.msaied.com/public/articles?search=llm)
- [streaming](https://www.msaied.com/public/articles?search=streaming)
- [agents](https://www.msaied.com/public/articles?search=agents)

 Frequently asked questions 
---------------------------

  Does streaming work with Laravel Octane?Yes, but you must use Swoole's chunked response or FrankenPHP's early-flush support rather than PHP's `ob\_flush`. Octane workers keep the connection open, so the SSE loop works — just replace `ob\_flush()/flush()` with the server-appropriate API and avoid storing streamed state on the worker.

   What happens if the model returns malformed JSON even with strict mode?With `strict: true` and a well-formed JSON Schema, the model is constrained by the API to match the schema. If the API itself returns an error or a refusal, the OpenAI PHP client throws an exception you can catch and retry. Always wrap the `json\_decode` call with `JSON\_THROW\_ON\_ERROR` as a final safety net.

   How do I track token usage per user for billing?The non-streaming response includes a `usage` object with `prompt\_tokens` and `completion\_tokens`. For streaming, request `stream\_options: \['include\_usage' =&gt; true\]` — the final SSE chunk will carry the usage data. Persist it in a `ai\_usage\_logs` table keyed by user and model for cost attribution.

   ![Mohamed Said](https://cdn.msaied.com/01M22N44A70A5MC2S599JP0MPH.webp)About the author
----------------

[Mohamed Said](https://www.msaied.com/public#person)Senior Backend Engineer specializing in Laravel, scalable SaaS platforms, APIs, and cloud infrastructure. I build secure, high-performance web applications that help businesses grow.

[About](https://www.msaied.com/public#about) [GitHub ↗](https://github.com/EG-Mohamed) [LinkedIn ↗](https://www.linkedin.com/in/msaiedm/) [WhatsApp ↗](https://wa.me/201094619204) [Email Address ↗](mailto:hello@msaied.com) [My CV ↗](https://drive.google.com/file/u/0/d/1MF20IPRJyzfy32mhEutjL5EpSls0w2Q8/view)  

   [Previous articleTyped PHP 8.3 Enums as Eloquent Casts, Route Parameters, and Validation Rules](https://www.msaied.com/public/articles/typed-php-83-enums-as-eloquent-casts-route-parameters-and-validation-rules) [Next articleFrankenPHP, OPcache JIT, and Preloading: Squeezing Real Throughput from Laravel](https://www.msaied.com/public/articles/frankenphp-opcache-jit-and-preloading-squeezing-real-throughput-from-laravel-3)  

   On this page
-------------

1. [The Problem With Naive LLM Integration](#the-problem-with-naive-llm-integration)
2. [Streaming Responses to the Browser](#streaming-responses-to-the-browser)
3. [Enforcing Token Budgets](#enforcing-token-budgets)
4. [1. Hard limit via max\_tokens](#1-hard-limit-via-codemax-tokenscode)
5. [2. Prompt token pre-check](#2-prompt-token-pre-check)
6. [Structured Output Contracts](#structured-output-contracts)
7. [Wiring It Together in a Job](#wiring-it-together-in-a-job)
8. [Key Takeaways](#key-takeaways)

 ###  Have a technical challenge?

 Tell me what you’re building. I reply within two working days.

[Start a conversation](https://www.msaied.com/public#contact) 

   Related articles
-----------------

 [ ![](https://cdn.msaied.com/740/cce86edc21eddcbdd2f2454fadaf9c70.png)  · 3 min read### The Pipeline Pattern in Laravel: Custom Pipelines Beyond Middleware

5 Oct 2026 ](https://www.msaied.com/public/articles/the-pipeline-pattern-in-laravel-custom-pipelines-beyond-middleware-1) [ ![](https://cdn.msaied.com/739/2d6897fdcdcf090613f96f72a64b8a78.png)  · 4 min read### MySQL Full-Text Search in Laravel: Indexes, Relevance Scoring, and Boolean Mode

4 Oct 2026 ](https://www.msaied.com/public/articles/mysql-full-text-search-in-laravel-indexes-relevance-scoring-and-boolean-mode) [ ![](https://cdn.msaied.com/738/073696a3fefe18bec825beec5ac658f5.png)  · 4 min read### Laravel Queue Rate-Limited Middleware: Throttling Jobs Without Losing Work

4 Oct 2026 ](https://www.msaied.com/public/articles/laravel-queue-rate-limited-middleware-throttling-jobs-without-losing-work) 

  Have a technical challenge?
----------------------------

Tell me what you’re building. I reply within two working days.

 [Discuss your project ↗](https://www.msaied.com/public#contact) 

  © 2026 Mohamed Said · Built with Laravel, meant to last.Senior Backend Engineer specializing in Laravel, scalable SaaS platforms, APIs, and cloud infrastructure. I build secure, high-performance web applications that help businesses grow.

 - [Home](https://www.msaied.com/public)
- [Articles](https://www.msaied.com/public/articles)
- [Certificates](https://www.msaied.com/public/certificates)
- [GitHub](https://github.com/EG-Mohamed)
- [LinkedIn](https://www.linkedin.com/in/msaiedm/)
- [WhatsApp](https://wa.me/201094619204)
- [Email Address](mailto:hello@msaied.com)
- [My CV](https://drive.google.com/file/u/0/d/1MF20IPRJyzfy32mhEutjL5EpSls0w2Q8/view)
- [Sitemap](https://www.msaied.com/public/sitemap.xml)
