Building Enterprise Node.js Microservices with PostgreSQL RAG and OpenAI Function Calling
Learn how to wire up Node.js microservices, PostgreSQL RAG pipelines, and OpenAI function calling for scalable AI‑augmented backends in 2026.

Why combine RAG and function calling in a microservice stack?
Retrieval‑augmented generation (RAG) gives you deterministic context from your own data, while OpenAI function calling turns LLM output into concrete API calls. When you wrap both inside a Node.js microservice, you get a system that can answer natural‑language queries and trigger business logic without a separate orchestration layer.
Service boundaries you actually need
In a typical enterprise you’ll have three logical services:
- Query Service: receives a user prompt, runs a vector search against PostgreSQL, and builds the RAG context.
- Orchestrator: forwards the prompt plus context to OpenAI, inspects the
function_callfield, and dispatches to the appropriate business service. - Domain Services: pure Node.js/Express or Fastify services that implement the functions (e.g.,
createInvoice,fetchCustomer).
Keeping these thin and versioned via OpenAPI means you can evolve each piece independently.
Setting up PostgreSQL for vector search
PostgreSQL 15+ ships with the pgvector extension. Install it once, then create a table that stores both the raw row and its embedding.
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE documents (
id uuid PRIMARY KEY,
content text NOT NULL,
embedding vector(1536) NOT NULL
);
CREATE INDEX ON documents USING ivfflat (embedding vector_cosine_ops) WITH (lists = 100);When a new document arrives, compute its embedding with OpenAI’s text-embedding-ada-002 and insert it. The query service then does a simple ORDER BY embedding <-> $1 LIMIT 5 to fetch the most relevant chunks.
OpenAI function calling workflow
The orchestrator sends a prompt like:
const response = await openai.chat.completions.create({
model: "gpt-4o-mini",
messages: [{ role: "user", content: userPrompt }],
functions: [
{ name: "createInvoice", description: "Create a new invoice", parameters: { type: "object", properties: { customerId: {type:"string"}, amount: {type:"number"} }, required: ["customerId","amount"] } }
]
});If the model decides a function is needed, it returns:
{"function_call": {"name": "createInvoice", "arguments": "{\"customerId\":\"c123\",\"amount\":250}"}}The orchestrator parses the JSON, calls the /invoices endpoint of the domain service, and returns the result to the client. This pattern eliminates brittle prompt‑parsing code and gives you an audit trail of every function invocation.
Putting it together in a Next.js API route
Here’s a minimal Next.js route that glues query, RAG, and function calling:
import { NextResponse } from 'next/server';
import { getRelevantDocs } from '@/lib/rag';
import { openai } from '@/lib/openai';
export async function POST(req: Request) {
const { prompt } = await req.json();
const docs = await getRelevantDocs(prompt); // vector search
const system = `Use the following context to answer:
${docs.map(d=>d.content).join('\n')}`;
const completion = await openai.chat.completions.create({
model: 'gpt-4o-mini',
messages: [{ role: 'system', content: system }, { role: 'user', content: prompt }],
functions: [{ name: 'createInvoice', ... }],
});
const { function_call } = completion.choices[0].message;
if (function_call) {
const args = JSON.parse(function_call.arguments);
const invoiceRes = await fetch('http://invoice-service/api/invoices', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify(args),
});
const result = await invoiceRes.json();
return NextResponse.json({ result });
}
return NextResponse.json({ answer: completion.choices[0].message.content });
}The route stays under 50 lines, yet you have a full RAG‑plus‑function pipeline.
Operational tips for 2026
- Version your function schemas. Store them in a Git‑tracked JSON file and bump a
schemaVersionfield when you add parameters. - Cache embeddings. Use Redis or a CDN edge cache to avoid recomputing embeddings on every request.
- Rate‑limit OpenAI calls. Even with function calling you can hit token limits; a simple token bucket per user works.
- Observability. Log the original prompt, retrieved docs, and the final function payload. Tools like OpenTelemetry + Loki make correlation easy.
By structuring your code this way you get a clean, testable backend that scales horizontally, respects data privacy (all context lives in PostgreSQL), and lets the LLM focus on what it does best: orchestrating your existing business logic.