RAG in Laravel with pgvector and Embeddings | Mohamed Said       [Skip to content](#main)  [ ![](https://cdn.msaied.com/01KT78WE565VEMM3PSNQAAB0MH.png) Mohamed SaidLaravel Backend Engineer ](https://www.msaied.com/public) - [Home](https://www.msaied.com/public)
- [Projects](https://www.msaied.com/public/projects)
- [Articles](https://www.msaied.com/public/articles)
- [Certificates](https://www.msaied.com/public/certificates)
- [About](https://www.msaied.com/public#about)

           [  Contact](https://www.msaied.com/public#contact) Menu 

Menu
----

Close 

 - [HomeStart here](https://www.msaied.com/public)
- [ProjectsCase studies](https://www.msaied.com/public/projects)
- [ArticlesEngineering notes](https://www.msaied.com/public/articles)
- [CertificatesCredentials](https://www.msaied.com/public/certificates)
- [AboutHow I work](https://www.msaied.com/public#about)
- [ContactGet in touch](https://www.msaied.com/public#contact)

  [Start a conversation](https://www.msaied.com/public#contact) [WhatsApp](https://wa.me/201094619204) [Email](mailto:hello@msaied.com) 

 1. [Home](https://www.msaied.com/public)
2. /
3. [Articles](https://www.msaied.com/public/articles)
4. /
5. Practical RAG in Laravel: pgvector, Embeddings, and Retrieval Pipelines

 Practical RAG in Laravel: pgvector, Embeddings, and Retrieval Pipelines
========================================================================

 Build a production-ready retrieval-augmented generation pipeline in Laravel using pgvector, OpenAI embeddings, and a clean retrieval abstraction — without reaching for a dedicated AI framework.

 ![](https://cdn.msaied.com/01M22N44A70A5MC2S599JP0MPH.webp) [Mohamed Said](https://www.msaied.com/public#person) Published 30 Jul 2026 · Updated 30 Jul 2026 · 4 min read

ShareCopy linkCopied

 ![Practical RAG in Laravel: pgvector, Embeddings, and Retrieval Pipelines](https://cdn.msaied.com/487/2db17c4f7927d4a229a4786c03e7b67b.png) 

  On this page +1. [Why RAG Instead of Fine-Tuning?](#why-rag-instead-of-fine-tuning)
2. [Setting Up pgvector](#setting-up-pgvector)
3. [Generating and Storing Embeddings](#generating-and-storing-embeddings)
4. [Retrieval: Nearest-Neighbour Query](#retrieval-nearest-neighbour-query)
5. [Building the Prompt](#building-the-prompt)
6. [Caching Embeddings](#caching-embeddings)
7. [Key Takeaways](#key-takeaways)

 Why RAG Instead of Fine-Tuning?
-------------------------------

Retrieval-augmented generation (RAG) lets you ground an LLM's answers in your own data without the cost and complexity of fine-tuning. The pattern is simple: embed your documents, store the vectors, retrieve the top-k nearest neighbours at query time, and inject them as context. Laravel's ecosystem — PostgreSQL, Eloquent, and the HTTP client — is more than enough to build this cleanly.

Setting Up pgvector
-------------------

Install the extension and add a migration:

```sql
CREATE EXTENSION IF NOT EXISTS vector;

```

```php
// database/migrations/2024_01_01_000000_add_embedding_to_documents.php
public function up(): void
{
    Schema::table('documents', function (Blueprint $table) {
        // 1536 dims for text-embedding-3-small
        $table->vector('embedding', 1536)->nullable();
    });

    DB::statement(
        'CREATE INDEX documents_embedding_hnsw_idx
         ON documents USING hnsw (embedding vector_cosine_ops)'
    );
}

```

Laravel's `Blueprint` doesn't know `vector` natively, so register a macro in a service provider:

```php
use Illuminate\Database\Schema\Blueprint;
use Illuminate\Support\Facades\Schema;

Blueprint::macro('vector', function (string $column, int $dimensions): \Illuminate\Database\Schema\ColumnDefinition {
    return $this->addColumn('vector', $column, compact('dimensions'));
});

// Register the type with Doctrine so migrations don't break
\Doctrine\DBAL\Types\Type::addType('vector', VectorType::class);

```

`VectorType` is a thin Doctrine type that serialises a PHP float array to the `[0.1,0.2,...]` string pgvector expects.

Generating and Storing Embeddings
---------------------------------

Wrap the OpenAI call in a dedicated action:

```php
final readonly class GenerateEmbedding
{
    public function __construct(private \Illuminate\Http\Client\Factory $http) {}

    /** @return float[] */
    public function handle(string $text): array
    {
        $response = $this->http
            ->withToken(config('services.openai.key'))
            ->post('https://api.openai.com/v1/embeddings', [
                'model' => 'text-embedding-3-small',
                'input' => $text,
            ])
            ->throw()
            ->json('data.0.embedding');

        return $response;
    }
}

```

Dispatch a job when a document is saved:

```php
final class EmbedDocument implements ShouldQueue
{
    use Dispatchable, Queueable;

    public function __construct(public readonly int $documentId) {}

    public function handle(GenerateEmbedding $action): void
    {
        $doc = Document::findOrFail($this->documentId);
        $vector = $action->handle($doc->body);

        // Store as pgvector literal
        DB::table('documents')
            ->where('id', $doc->id)
            ->update(['embedding' => '[' . implode(',', $vector) . ']']);
    }
}

```

Retrieval: Nearest-Neighbour Query
----------------------------------

A clean retrieval abstraction keeps the pgvector SQL out of your controllers:

```php
final readonly class DocumentRetriever
{
    public function __construct(
        private GenerateEmbedding $embedder,
        private int $topK = 5,
    ) {}

    /** @return \Illuminate\Support\Collection */
    public function retrieve(string $query): \Illuminate\Support\Collection
    {
        $vector = '[' . implode(',', $this->embedder->handle($query)) . ']';

        return Document::query()
            ->selectRaw('*, embedding  ? AS distance', [$vector])
            ->whereNotNull('embedding')
            ->orderBy('distance')
            ->limit($this->topK)
            ->get();
    }
}

```

The `` operator is pgvector's cosine distance. For inner-product similarity use ``.

Building the Prompt
-------------------

```php
final readonly class RagPipeline
{
    public function __construct(
        private DocumentRetriever $retriever,
        private \Illuminate\Http\Client\Factory $http,
    ) {}

    public function answer(string $question): string
    {
        $context = $this->retriever->retrieve($question)
            ->map(fn (Document $d) => "- {$d->title}: {$d->body}")
            ->implode("\n");

        return $this->http
            ->withToken(config('services.openai.key'))
            ->post('https://api.openai.com/v1/chat/completions', [
                'model' => 'gpt-4o-mini',
                'messages' => [
                    ['role' => 'system', 'content' => "Answer using only the context below.\n\n{$context}"],
                    ['role' => 'user',   'content' => $question],
                ],
            ])
            ->throw()
            ->json('choices.0.message.content');
    }
}

```

### Caching Embeddings

Embedding the same query repeatedly wastes tokens. Cache by hash:

```php
$cacheKey = 'embedding:' . hash('xxh128', $text);
$vector = Cache::remember($cacheKey, now()->addDay(), fn () => $action->handle($text));

```

Key Takeaways
-------------

- Register a `Blueprint::macro` for `vector` columns; pair it with a Doctrine type for migration compatibility.
- Use HNSW indexes (`vector_cosine_ops`) for sub-millisecond ANN at scale — IVFFlat requires a `VACUUM` before it becomes useful.
- Keep embedding generation in a queued job; retrieval at request time is fast enough for synchronous use.
- Wrap retrieval behind a `DocumentRetriever` class so you can swap pgvector for another store without touching controllers.
- Cache query embeddings by content hash to avoid redundant API calls on repeated questions.

- [laravel](https://www.msaied.com/public/articles?search=laravel)
- [ai](https://www.msaied.com/public/articles?search=ai)
- [pgvector](https://www.msaied.com/public/articles?search=pgvector)
- [postgresql](https://www.msaied.com/public/articles?search=postgresql)
- [rag](https://www.msaied.com/public/articles?search=rag)

 Frequently asked questions 
---------------------------

  Do I need a dedicated vector database, or is pgvector enough?pgvector with an HNSW index handles millions of vectors comfortably for most SaaS workloads. A dedicated store like Qdrant or Pinecone adds operational overhead that is rarely justified until you exceed tens of millions of vectors or need advanced filtering that pgvector cannot express efficiently.

   How do I handle documents longer than the embedding model's token limit?Chunk the document before embedding — typically 512–1024 tokens with a small overlap (e.g. 64 tokens) to preserve context across chunk boundaries. Store each chunk as a separate row with a foreign key back to the parent document, then deduplicate retrieved chunks by document at query time.

   Can I test the retrieval pipeline without hitting the OpenAI API?Yes. Bind a fake implementation of GenerateEmbedding in your test service provider that returns a deterministic float array. Insert pre-computed embeddings into the test database and assert that DocumentRetriever returns the expected rows — no HTTP calls required.

   ![Mohamed Said](https://cdn.msaied.com/01M22N44A70A5MC2S599JP0MPH.webp)About the author
----------------

[Mohamed Said](https://www.msaied.com/public#person)Senior Backend Engineer specializing in Laravel, scalable SaaS platforms, APIs, and cloud infrastructure. I build secure, high-performance web applications that help businesses grow.

[About](https://www.msaied.com/public#about) [GitHub ↗](https://github.com/EG-Mohamed) [LinkedIn ↗](https://www.linkedin.com/in/msaiedm/) [WhatsApp ↗](https://wa.me/201094619204) [Email Address ↗](mailto:hello@msaied.com) [My CV ↗](https://drive.google.com/file/u/0/d/1MF20IPRJyzfy32mhEutjL5EpSls0w2Q8/view)  

   [Previous articleLaravel Octane + FrankenPHP: Persistent Services, Boot-Once Singletons, and Safe State](https://www.msaied.com/public/articles/laravel-octane-frankenphp-persistent-services-boot-once-singletons-and-safe-state) [Next articleLaravel Observers vs. Model Events: Choosing the Right Hook for Domain Side-Effects](https://www.msaied.com/public/articles/laravel-observers-vs-model-events-choosing-the-right-hook-for-domain-side-effects)  

   On this page
-------------

1. [Why RAG Instead of Fine-Tuning?](#why-rag-instead-of-fine-tuning)
2. [Setting Up pgvector](#setting-up-pgvector)
3. [Generating and Storing Embeddings](#generating-and-storing-embeddings)
4. [Retrieval: Nearest-Neighbour Query](#retrieval-nearest-neighbour-query)
5. [Building the Prompt](#building-the-prompt)
6. [Caching Embeddings](#caching-embeddings)
7. [Key Takeaways](#key-takeaways)

 ###  Have a technical challenge?

 Tell me what you’re building. I reply within two working days.

[Start a conversation](https://www.msaied.com/public#contact) 

   Related articles
-----------------

 [ ![](https://cdn.msaied.com/740/cce86edc21eddcbdd2f2454fadaf9c70.png)  · 3 min read### The Pipeline Pattern in Laravel: Custom Pipelines Beyond Middleware

5 Oct 2026 ](https://www.msaied.com/public/articles/the-pipeline-pattern-in-laravel-custom-pipelines-beyond-middleware-1) [ ![](https://cdn.msaied.com/739/2d6897fdcdcf090613f96f72a64b8a78.png)  · 4 min read### MySQL Full-Text Search in Laravel: Indexes, Relevance Scoring, and Boolean Mode

4 Oct 2026 ](https://www.msaied.com/public/articles/mysql-full-text-search-in-laravel-indexes-relevance-scoring-and-boolean-mode) [ ![](https://cdn.msaied.com/738/073696a3fefe18bec825beec5ac658f5.png)  · 4 min read### Laravel Queue Rate-Limited Middleware: Throttling Jobs Without Losing Work

4 Oct 2026 ](https://www.msaied.com/public/articles/laravel-queue-rate-limited-middleware-throttling-jobs-without-losing-work) 

  Have a technical challenge?
----------------------------

Tell me what you’re building. I reply within two working days.

 [Discuss your project ↗](https://www.msaied.com/public#contact) 

  © 2026 Mohamed Said · Built with Laravel, meant to last.Senior Backend Engineer specializing in Laravel, scalable SaaS platforms, APIs, and cloud infrastructure. I build secure, high-performance web applications that help businesses grow.

 - [Home](https://www.msaied.com/public)
- [Articles](https://www.msaied.com/public/articles)
- [Certificates](https://www.msaied.com/public/certificates)
- [GitHub](https://github.com/EG-Mohamed)
- [LinkedIn](https://www.linkedin.com/in/msaiedm/)
- [WhatsApp](https://wa.me/201094619204)
- [Email Address](mailto:hello@msaied.com)
- [My CV](https://drive.google.com/file/u/0/d/1MF20IPRJyzfy32mhEutjL5EpSls0w2Q8/view)
- [Sitemap](https://www.msaied.com/public/sitemap.xml)
