Skip to content

RAG Q&A endpoint ​

Build a minimal retrieval-augmented Q&A flow with @pondoknusa/rag and @pondoknusa/vector.

Scaffold ​

bash
pondoknusa new knowledge-base --ai
npm install
pondoknusa vector:install

The --ai flag adds vector config, embed jobs, models, and example routes.

Ingest documents ​

Prefer inline content on HTTP endpoints. For filesystem ingest from trusted jobs/CLI, pass rootDir so paths cannot escape the ingest root:

typescript
import { ingestDocument, ingestFile } from '@pondoknusa/rag';
import { Document } from './models/Document.js';

await ingestDocument(Document, {
  source: 'handbook',
  content: '# Handbook\n...',
});

await ingestFile(Document, 'handbook.pdf', {
  source: 'handbook',
  rootDir: 'storage/documents',
  chunkSize: 800,
});

Embed chunks:

bash
pondoknusa vector:embed --model=Document

Ask endpoint ​

typescript
import { Route } from '@pondoknusa/core';
import { Response } from '@pondoknusa/http';
import { Rag } from '@pondoknusa/rag';
import { inferenceChat, inferenceEmbed } from '@pondoknusa/inference';
import { Document } from './models/Document.js';

const rag = new Rag({
  model: Document,
  embed: async (text) => (await inferenceEmbed(text))[0]!,
});

Route.post('/api/ask', async (request) => {
  const { question } = await request.json<{ question: string }>();
  const chunks = await rag.retrieve(question, { topK: 5 });
  const prompt = rag.buildPrompt(question, chunks);

  const { content } = await inferenceChat([{ role: 'user', content: prompt }]);
  return Response.json({ answer: content });
});

Use your preferred LLM SDK in the app layer — Pondoknusa handles storage, retrieval, and prompt templates.

Example app ​

See examples/rag for ingest → embed → ask → stream with GraphQL read API.

Released under the MIT License.