Streaming AI in Laravel: Token Budgets &amp; Structured Output | Mohamed Said       [Skip to content](#main)  [ ![](https://cdn.msaied.com/01KT78WE565VEMM3PSNQAAB0MH.png) Mohamed SaidLaravel Backend Engineer ](https://www.msaied.com/public) - [Home](https://www.msaied.com/public)
- [Projects](https://www.msaied.com/public/projects)
- [Articles](https://www.msaied.com/public/articles)
- [Certificates](https://www.msaied.com/public/certificates)
- [About](https://www.msaied.com/public#about)

           [  Contact](https://www.msaied.com/public#contact) Menu 

Menu
----

Close 

 - [HomeStart here](https://www.msaied.com/public)
- [ProjectsCase studies](https://www.msaied.com/public/projects)
- [ArticlesEngineering notes](https://www.msaied.com/public/articles)
- [CertificatesCredentials](https://www.msaied.com/public/certificates)
- [AboutHow I work](https://www.msaied.com/public#about)
- [ContactGet in touch](https://www.msaied.com/public#contact)

  [Start a conversation](https://www.msaied.com/public#contact) [WhatsApp](https://wa.me/201094619204) [Email](mailto:hello@msaied.com) 

 1. [Home](https://www.msaied.com/public)
2. /
3. [Articles](https://www.msaied.com/public/articles)
4. /
5. Streaming AI Responses in Laravel: Token Budgets, Structured Output, and Production Contracts

 Streaming AI Responses in Laravel: Token Budgets, Structured Output, and Production Contracts
==============================================================================================

 Learn how to stream AI completions in Laravel with hard token budgets, enforce structured JSON output contracts, and handle partial failures gracefully in production — without leaking state across Octane workers.

 ![](https://cdn.msaied.com/01M22N44A70A5MC2S599JP0MPH.webp) [Mohamed Said](https://www.msaied.com/public#person) Published 15 Jun 2026 · Updated 15 Jun 2026 · 4 min read

ShareCopy linkCopied

 ![Streaming AI Responses in Laravel: Token Budgets, Structured Output, and Production Contracts](https://cdn.msaied.com/199/0042e3985f7cb3bfdaeb2d1c76c79489.png) 

  On this page +1. [The Problem With Naive AI Integration](#the-problem-with-naive-ai-integration)
2. [1. Streaming Completions Without Blocking the Worker](#1-streaming-completions-without-blocking-the-worker)
3. [2. Enforcing Token Budgets](#2-enforcing-token-budgets)
4. [3. Structured Output Contracts with JSON Schema](#3-structured-output-contracts-with-json-schema)
5. [4. Octane Safety: No Static State, No Singleton Leakage](#4-octane-safety-no-static-state-no-singleton-leakage)
6. [Takeaways](#takeaways)

 The Problem With Naive AI Integration
-------------------------------------

Most Laravel AI tutorials end at `Http::post()` and `json_decode()`. In production you face three harder problems: responses that arrive token-by-token (streaming), models that hallucinate structure (unstructured output), and long-lived Octane workers that silently carry state between requests. This article tackles all three with concrete, opinionated patterns.

---

1. Streaming Completions Without Blocking the Worker
----------------------------------------------------

OpenAI's streaming API sends `text/event-stream` chunks. Laravel's HTTP client wraps Guzzle, so you can consume the stream lazily:

```php
use Illuminate\Support\Facades\Http;

function streamCompletion(string $prompt): \Generator
{
    $response = Http::withToken(config('services.openai.key'))
        ->withOptions(['stream' => true])
        ->post('https://api.openai.com/v1/chat/completions', [
            'model'      => 'gpt-4o',
            'stream'     => true,
            'max_tokens' => 512,
            'messages'   => [['role' => 'user', 'content' => $prompt]],
        ]);

    $body = $response->toPsrResponse()->getBody();

    while (! $body->eof()) {
        $line = trim($body->read(4096));
        if (str_starts_with($line, 'data: ') && $line !== 'data: [DONE]') {
            $chunk = json_decode(substr($line, 6), true);
            yield $chunk['choices'][0]['delta']['content'] ?? '';
        }
    }
}

```

Return this generator from a `StreamedResponse` so Nginx flushes each chunk immediately:

```php
Route::get('/stream', function () {
    return response()->stream(function () {
        foreach (streamCompletion('Explain CQRS in one paragraph') as $token) {
            echo "data: {$token}\n\n";
            ob_flush();
            flush();
        }
    }, 200, ['Content-Type' => 'text/event-stream', 'X-Accel-Buffering' => 'no']);
});

```

`X-Accel-Buffering: no` is mandatory when Nginx sits in front — without it the proxy buffers the entire response.

---

2. Enforcing Token Budgets
--------------------------

`max_tokens` is a ceiling, not a guarantee. A budget-aware wrapper counts tokens before the call and aborts early if the prompt itself is too large:

```php
final class TokenBudget
{
    public function __construct(
        private readonly int $maxPromptTokens = 3_000,
        private readonly int $maxCompletionTokens = 512,
    ) {}

    /** Rough estimate: 1 token ≈ 4 chars for English prose */
    public function promptFits(string $prompt): bool
    {
        return (int) ceil(mb_strlen($prompt) / 4) maxPromptTokens;
    }

    public function completionLimit(): int
    {
        return $this->maxCompletionTokens;
    }
}

```

Bind it as a singleton in `AppServiceProvider` and inject it wherever you build prompts. This prevents runaway costs when user-supplied context is large.

---

3. Structured Output Contracts with JSON Schema
-----------------------------------------------

OpenAI's `response_format` with `json_schema` mode guarantees the model returns valid JSON matching your schema — or it refuses rather than hallucinating:

```php
$schema = [
    'type'       => 'object',
    'properties' => [
        'summary'    => ['type' => 'string'],
        'confidence' => ['type' => 'number', 'minimum' => 0, 'maximum' => 1],
        'tags'       => ['type' => 'array', 'items' => ['type' => 'string']],
    ],
    'required'             => ['summary', 'confidence', 'tags'],
    'additionalProperties' => false,
];

$result = Http::withToken(config('services.openai.key'))
    ->post('https://api.openai.com/v1/chat/completions', [
        'model'           => 'gpt-4o-2024-08-06',
        'max_tokens'      => 256,
        'response_format' => [
            'type'        => 'json_schema',
            'json_schema' => ['name' => 'analysis', 'strict' => true, 'schema' => $schema],
        ],
        'messages' => [['role' => 'user', 'content' => "Analyse: {$text}"]],
    ])->json('choices.0.message.content');

$dto = AnalysisResult::fromArray(json_decode($result, true));

```

Map the validated JSON straight into a typed DTO — no defensive `isset()` chains needed.

---

4. Octane Safety: No Static State, No Singleton Leakage
-------------------------------------------------------

Octane workers are long-lived. Any static property or singleton that accumulates per-request data will bleed across users. For AI work:

- **Never** store conversation history in a singleton. Use the session or a database-backed `Conversation` model.
- Bind AI client wrappers as `scoped()` (reset per request) rather than `singleton()`.
- Use `defer()` for logging token usage so it runs after the response is sent.

```php
// AppServiceProvider
$this->app->scoped(AiClient::class, fn () => new AiClient(
    apiKey: config('services.openai.key'),
));

```

---

Takeaways
---------

- Stream via `Http::withOptions(['stream' => true])` and yield chunks through a `StreamedResponse`.
- Set `X-Accel-Buffering: no` when Nginx proxies the stream.
- Estimate prompt token size before the API call to enforce hard cost budgets.
- Use OpenAI's `json_schema` response format to get guaranteed-valid structured output.
- Register AI clients as `scoped()` bindings in Octane to prevent cross-request state leakage.
- Map structured responses directly into typed DTOs — skip defensive null-checking.

- [laravel](https://www.msaied.com/public/articles?search=laravel)
- [ai](https://www.msaied.com/public/articles?search=ai)
- [streaming](https://www.msaied.com/public/articles?search=streaming)
- [php](https://www.msaied.com/public/articles?search=php)

 Frequently asked questions 
---------------------------

  Does Laravel's HTTP client support streaming responses natively?Yes. Pass `\['stream' =&gt; true\]` in `withOptions()` and call `toPsrResponse()-&gt;getBody()` to get a PSR-7 stream you can read incrementally. Wrap the output in `response()-&gt;stream()` to flush tokens to the browser as they arrive.

   What is the difference between `max\_tokens` and a token budget?`max\_tokens` caps the model's completion length on the API side. A token budget is an application-level guard that estimates prompt size before the call and rejects requests that would exceed your cost or context-window limits — preventing expensive or failed API calls entirely.

   How do I prevent AI client state from leaking between Octane requests?Register your AI client wrapper with `$this-&gt;app-&gt;scoped()` instead of `singleton()`. Scoped bindings are flushed and rebuilt at the start of each request cycle, so no per-request data (conversation history, accumulated tokens) survives into the next request.

   ![Mohamed Said](https://cdn.msaied.com/01M22N44A70A5MC2S599JP0MPH.webp)About the author
----------------

[Mohamed Said](https://www.msaied.com/public#person)Senior Backend Engineer specializing in Laravel, scalable SaaS platforms, APIs, and cloud infrastructure. I build secure, high-performance web applications that help businesses grow.

[About](https://www.msaied.com/public#about) [GitHub ↗](https://github.com/EG-Mohamed) [LinkedIn ↗](https://www.linkedin.com/in/msaiedm/) [WhatsApp ↗](https://wa.me/201094619204) [Email Address ↗](mailto:hello@msaied.com) [My CV ↗](https://drive.google.com/file/u/0/d/1MF20IPRJyzfy32mhEutjL5EpSls0w2Q8/view)  

   [Previous articleFilament v3 to v4 Migration: Breaking Changes and Practical Refactor Patterns](https://www.msaied.com/public/articles/filament-v3-to-v4-migration-breaking-changes-and-practical-refactor-patterns) [Next articleMulti-Tenant SaaS with Laravel: Scoping Queries, Resolving Tenants, and Avoiding Data Leaks](https://www.msaied.com/public/articles/multi-tenant-saas-with-laravel-scoping-queries-resolving-tenants-and-avoiding-data-leaks)  

   On this page
-------------

1. [The Problem With Naive AI Integration](#the-problem-with-naive-ai-integration)
2. [1. Streaming Completions Without Blocking the Worker](#1-streaming-completions-without-blocking-the-worker)
3. [2. Enforcing Token Budgets](#2-enforcing-token-budgets)
4. [3. Structured Output Contracts with JSON Schema](#3-structured-output-contracts-with-json-schema)
5. [4. Octane Safety: No Static State, No Singleton Leakage](#4-octane-safety-no-static-state-no-singleton-leakage)
6. [Takeaways](#takeaways)

 ###  Have a technical challenge?

 Tell me what you’re building. I reply within two working days.

[Start a conversation](https://www.msaied.com/public#contact) 

   Related articles
-----------------

 [ ![](https://cdn.msaied.com/740/cce86edc21eddcbdd2f2454fadaf9c70.png)  · 3 min read### The Pipeline Pattern in Laravel: Custom Pipelines Beyond Middleware

5 Oct 2026 ](https://www.msaied.com/public/articles/the-pipeline-pattern-in-laravel-custom-pipelines-beyond-middleware-1) [ ![](https://cdn.msaied.com/739/2d6897fdcdcf090613f96f72a64b8a78.png)  · 4 min read### MySQL Full-Text Search in Laravel: Indexes, Relevance Scoring, and Boolean Mode

4 Oct 2026 ](https://www.msaied.com/public/articles/mysql-full-text-search-in-laravel-indexes-relevance-scoring-and-boolean-mode) [ ![](https://cdn.msaied.com/738/073696a3fefe18bec825beec5ac658f5.png)  · 4 min read### Laravel Queue Rate-Limited Middleware: Throttling Jobs Without Losing Work

4 Oct 2026 ](https://www.msaied.com/public/articles/laravel-queue-rate-limited-middleware-throttling-jobs-without-losing-work) 

  Have a technical challenge?
----------------------------

Tell me what you’re building. I reply within two working days.

 [Discuss your project ↗](https://www.msaied.com/public#contact) 

  © 2026 Mohamed Said · Built with Laravel, meant to last.Senior Backend Engineer specializing in Laravel, scalable SaaS platforms, APIs, and cloud infrastructure. I build secure, high-performance web applications that help businesses grow.

 - [Home](https://www.msaied.com/public)
- [Articles](https://www.msaied.com/public/articles)
- [Certificates](https://www.msaied.com/public/certificates)
- [GitHub](https://github.com/EG-Mohamed)
- [LinkedIn](https://www.linkedin.com/in/msaiedm/)
- [WhatsApp](https://wa.me/201094619204)
- [Email Address](mailto:hello@msaied.com)
- [My CV](https://drive.google.com/file/u/0/d/1MF20IPRJyzfy32mhEutjL5EpSls0w2Q8/view)
- [Sitemap](https://www.msaied.com/public/sitemap.xml)
