Laravel AI Agents: Streaming, Token Budgets &amp; Typed Output | Mohamed Said       [Skip to content](#main)  [ ![](https://cdn.msaied.com/01KT78WE565VEMM3PSNQAAB0MH.png) Mohamed SaidLaravel Backend Engineer ](https://www.msaied.com/public) - [Home](https://www.msaied.com/public)
- [Projects](https://www.msaied.com/public/projects)
- [Articles](https://www.msaied.com/public/articles)
- [Certificates](https://www.msaied.com/public/certificates)
- [About](https://www.msaied.com/public#about)

           [  Contact](https://www.msaied.com/public#contact) Menu 

Menu
----

Close 

 - [HomeStart here](https://www.msaied.com/public)
- [ProjectsCase studies](https://www.msaied.com/public/projects)
- [ArticlesEngineering notes](https://www.msaied.com/public/articles)
- [CertificatesCredentials](https://www.msaied.com/public/certificates)
- [AboutHow I work](https://www.msaied.com/public#about)
- [ContactGet in touch](https://www.msaied.com/public#contact)

  [Start a conversation](https://www.msaied.com/public#contact) [WhatsApp](https://wa.me/201094619204) [Email](mailto:hello@msaied.com) 

 1. [Home](https://www.msaied.com/public)
2. /
3. [Articles](https://www.msaied.com/public/articles)
4. /
5. Production AI Agents in Laravel: Streaming, Token Budgets, and Structured Output Contracts

 Production AI Agents in Laravel: Streaming, Token Budgets, and Structured Output Contracts
===========================================================================================

 Ship reliable AI agents in Laravel by combining streamed responses, hard token budgets, and typed structured-output contracts — without letting LLM non-determinism bleed into your domain layer.

 ![](https://cdn.msaied.com/01M22N44A70A5MC2S599JP0MPH.webp) [Mohamed Said](https://www.msaied.com/public#person) Published 22 Jun 2026 · Updated 22 Jun 2026 · 3 min read

ShareCopy linkCopied

 ![Production AI Agents in Laravel: Streaming, Token Budgets, and Structured Output Contracts](https://cdn.msaied.com/262/4220ef4ee3875717c6384f298d891a3f.png) 

  On this page +1. [The Problem With Naive AI Integration](#the-problem-with-naive-ai-integration)
2. [Streaming Responses to the Browser](#streaming-responses-to-the-browser)
3. [Token Budget Guard](#token-budget-guard)
4. [Structured Output Contracts](#structured-output-contracts)
5. [Keeping the Domain Clean](#keeping-the-domain-clean)
6. [Takeaways](#takeaways)

 The Problem With Naive AI Integration
-------------------------------------

Most Laravel + LLM tutorials stop at `Http::post('https://api.openai.com/v1/chat/completions', [...])` and call it done. That works for demos. In production you need three things the tutorials skip:

1. **Streaming** — users shouldn't stare at a spinner for 8 seconds.
2. **Token budgets** — unbounded prompts destroy your billing and latency SLAs.
3. **Structured output contracts** — raw JSON strings from an LLM are not domain objects.

Let's solve all three without a heavy third-party SDK.

---

Streaming Responses to the Browser
----------------------------------

OpenAI's `stream: true` returns server-sent events. Laravel's `StreamedResponse` pipes them straight to the client.

```php
// app/Http/Controllers/AgentController.php
public function stream(Request $request): StreamedResponse
{
    $prompt = $request->validated()['prompt'];

    return response()->stream(function () use ($prompt) {
        $stream = Http::withToken(config('services.openai.key'))
            ->withOptions(['stream' => true])
            ->post('https://api.openai.com/v1/chat/completions', [
                'model'    => 'gpt-4o-mini',
                'stream'   => true,
                'messages' => [['role' => 'user', 'content' => $prompt]],
            ])->toPsrResponse()->getBody();

        while (! $stream->eof()) {
            $line = trim($stream->read(512));
            if (str_starts_with($line, 'data: ')) {
                $payload = substr($line, 6);
                if ($payload === '[DONE]') break;
                $delta = json_decode($payload, true)['choices'][0]['delta']['content'] ?? '';
                echo "data: {$delta}\n\n";
                ob_flush(); flush();
            }
        }
    }, 200, ['Content-Type' => 'text/event-stream', 'X-Accel-Buffering' => 'no']);
}

```

The `X-Accel-Buffering: no` header is essential when Nginx sits in front — without it, Nginx buffers the whole response.

---

Token Budget Guard
------------------

Never let user input dictate prompt size. Enforce a budget before the HTTP call.

```php
// app/AI/TokenBudget.php
final class TokenBudget
{
    private const CHARS_PER_TOKEN = 4; // rough heuristic
    private const MAX_INPUT_TOKENS = 1_500;

    public static function enforce(string $text): string
    {
        $limit = self::MAX_INPUT_TOKENS * self::CHARS_PER_TOKEN;

        if (strlen($text) post('https://api.openai.com/v1/chat/completions', [
                'model'           => 'gpt-4o-mini',
                'max_tokens'      => 256,
                'response_format' => [
                    'type'        => 'json_schema',
                    'json_schema' => [
                        'name'   => 'sentiment_result',
                        'strict' => true,
                        'schema' => [
                            'type'       => 'object',
                            'properties' => [
                                'label'  => ['type' => 'string', 'enum' => ['positive','neutral','negative']],
                                'score'  => ['type' => 'number'],
                                'reason' => ['type' => 'string'],
                            ],
                            'required'            => ['label','score','reason'],
                            'additionalProperties'=> false,
                        ],
                    ],
                ],
                'messages' => [
                    ['role' => 'system', 'content' => 'Analyse sentiment. Reply only with the JSON schema.'],
                    ['role' => 'user',   'content' => $text],
                ],
            ])->throw()->json();

        $raw = json_decode(
            $response['choices'][0]['message']['content'],
            true,
            flags: JSON_THROW_ON_ERROR
        );

        return SentimentResult::fromArray($raw);
    }
}

```

With `strict: true` the model is constrained to the schema at the API level — you still validate on your side, but you'll rarely see a mismatch.

---

Keeping the Domain Clean
------------------------

The `SentimentAgent` returns a typed DTO. Nothing in your domain layer touches raw LLM strings. If you swap providers tomorrow, only the agent changes — every caller keeps working.

Wrap the agent in a queued job for non-interactive workloads, and inject it via the service container so Pest can swap in a fake:

```php
// tests/Feature/SentimentTest.php
it('classifies positive text', function () {
    $this->instance(SentimentAgent::class, new class {
        public function analyse(string $text): SentimentResult {
            return new SentimentResult('positive', 0.95, 'stub');
        }
    });

    $result = app(SentimentAgent::class)->analyse('Great product!');
    expect($result->label)->toBe('positive');
});

```

---

Takeaways
---------

- Stream via `response()->stream()` and set `X-Accel-Buffering: no` for Nginx.
- Enforce token budgets *before* the HTTP call, not after.
- Use `response_format.json_schema` with `strict: true` to pin model output shape.
- Map LLM JSON to typed readonly DTOs immediately — keep raw strings out of your domain.
- Inject agents through the container so tests can swap fakes without HTTP calls.

- [laravel](https://www.msaied.com/public/articles?search=laravel)
- [ai](https://www.msaied.com/public/articles?search=ai)
- [llm](https://www.msaied.com/public/articles?search=llm)
- [php](https://www.msaied.com/public/articles?search=php)

 Frequently asked questions 
---------------------------

  Does `response\_format: json\_schema` work with all OpenAI models?No. As of the current API, structured output with `strict: true` is supported on `gpt-4o`, `gpt-4o-mini`, and later snapshots. Older models like `gpt-3.5-turbo` support `json\_object` mode only, which does not enforce a schema.

   How do I handle streaming in a queued job rather than an HTTP response?In a job you don't need SSE. Disable streaming (`stream: false`), collect the full completion, then persist or broadcast the result. Streaming is only valuable when a human is waiting in real time.

   Is the 4-characters-per-token heuristic accurate enough for production?It's a conservative approximation for English text. For precise budgeting, use a tokeniser library such as `yethee/tiktoken` which implements the actual BPE encoding. The heuristic is fine as a cheap pre-flight guard.

   ![Mohamed Said](https://cdn.msaied.com/01M22N44A70A5MC2S599JP0MPH.webp)About the author
----------------

[Mohamed Said](https://www.msaied.com/public#person)Senior Backend Engineer specializing in Laravel, scalable SaaS platforms, APIs, and cloud infrastructure. I build secure, high-performance web applications that help businesses grow.

[About](https://www.msaied.com/public#about) [GitHub ↗](https://github.com/EG-Mohamed) [LinkedIn ↗](https://www.linkedin.com/in/msaiedm/) [WhatsApp ↗](https://wa.me/201094619204) [Email Address ↗](mailto:hello@msaied.com) [My CV ↗](https://drive.google.com/file/u/0/d/1MF20IPRJyzfy32mhEutjL5EpSls0w2Q8/view)  

   [Previous articleFilament v3 to v4: Breaking Changes and Practical Refactor Patterns](https://www.msaied.com/public/articles/filament-v3-to-v4-breaking-changes-and-practical-refactor-patterns) [Next articleMulti-Tenant SaaS in Laravel: Isolating Tenant State with Scoped Singletons](https://www.msaied.com/public/articles/multi-tenant-saas-in-laravel-isolating-tenant-state-with-scoped-singletons)  

   On this page
-------------

1. [The Problem With Naive AI Integration](#the-problem-with-naive-ai-integration)
2. [Streaming Responses to the Browser](#streaming-responses-to-the-browser)
3. [Token Budget Guard](#token-budget-guard)
4. [Structured Output Contracts](#structured-output-contracts)
5. [Keeping the Domain Clean](#keeping-the-domain-clean)
6. [Takeaways](#takeaways)

 ###  Have a technical challenge?

 Tell me what you’re building. I reply within two working days.

[Start a conversation](https://www.msaied.com/public#contact) 

   Related articles
-----------------

 [ ![](https://cdn.msaied.com/740/cce86edc21eddcbdd2f2454fadaf9c70.png)  · 3 min read### The Pipeline Pattern in Laravel: Custom Pipelines Beyond Middleware

5 Oct 2026 ](https://www.msaied.com/public/articles/the-pipeline-pattern-in-laravel-custom-pipelines-beyond-middleware-1) [ ![](https://cdn.msaied.com/739/2d6897fdcdcf090613f96f72a64b8a78.png)  · 4 min read### MySQL Full-Text Search in Laravel: Indexes, Relevance Scoring, and Boolean Mode

4 Oct 2026 ](https://www.msaied.com/public/articles/mysql-full-text-search-in-laravel-indexes-relevance-scoring-and-boolean-mode) [ ![](https://cdn.msaied.com/738/073696a3fefe18bec825beec5ac658f5.png)  · 4 min read### Laravel Queue Rate-Limited Middleware: Throttling Jobs Without Losing Work

4 Oct 2026 ](https://www.msaied.com/public/articles/laravel-queue-rate-limited-middleware-throttling-jobs-without-losing-work) 

  Have a technical challenge?
----------------------------

Tell me what you’re building. I reply within two working days.

 [Discuss your project ↗](https://www.msaied.com/public#contact) 

  © 2026 Mohamed Said · Built with Laravel, meant to last.Senior Backend Engineer specializing in Laravel, scalable SaaS platforms, APIs, and cloud infrastructure. I build secure, high-performance web applications that help businesses grow.

 - [Home](https://www.msaied.com/public)
- [Articles](https://www.msaied.com/public/articles)
- [Certificates](https://www.msaied.com/public/certificates)
- [GitHub](https://github.com/EG-Mohamed)
- [LinkedIn](https://www.linkedin.com/in/msaiedm/)
- [WhatsApp](https://wa.me/201094619204)
- [Email Address](mailto:hello@msaied.com)
- [My CV](https://drive.google.com/file/u/0/d/1MF20IPRJyzfy32mhEutjL5EpSls0w2Q8/view)
- [Sitemap](https://www.msaied.com/public/sitemap.xml)
