
#Agora — RAG System for AI Governance Documents
Agora is a production-grade Retrieval-Augmented Generation system built for the ETO AGORA Corpus — a collection of AI governance documents, regulations, and laws from jurisdictions worldwide. It lets users ask natural-language questions about policy documents and get grounded, source-attributed answers instead of hallucinated summaries.
#What It Does
Agora answers questions over AI governance documentation with an emphasis on accuracy over speed. Every response is anchored in retrieved documents, source-attributed, and filtered through a governance-analyst persona that distinguishes between what documents prohibit, require, recommend, or permit.
#Key Features
Intelligent Query Processing
- Sub-query decomposition splits multi-part questions into 1–5 focused retrieval passes for comprehensive coverage
- Conversational intent detection short-circuits RAG for greetings and pleasantries
- Parallel embedding and retrieval across sub-queries
Context Management
- Multi-turn conversation memory via Upstash Redis (30-minute model context window, 7-day UI history)
- Automatic deduplication of repeated content
- Source citations with cosine similarity scores
Production Infrastructure
- Document ingestion pipeline supporting
.txtand PDF formats - Background task management for large uploads
- Namespace isolation for multiple document collections
#Tech Stack
| Component | Technology |
|---|---|
| Backend | FastAPI + Uvicorn |
| Embeddings | Gemini Embedding 2 (1536D) |
| Text Generation | Gemini 2.5 Flash |
| Vector Store | Pinecone (serverless, cosine) |
| Session Memory | Upstash Redis |
| Frontend | Streamlit |
| PDF Parsing | pdfplumber |
#Solution Architecture
Request pipeline — every question runs through five stages:
User Question
→ Sub-Query Decomposition (Gemini 2.5 Flash classifies intent: conversational / single / multi-part)
→ Parallel Embedding + Pinecone Retrieval (top-4 chunks per sub-query, cosine similarity)
→ Context Assembly (dedupe by chunk-text prefix, rank by score, pull last 5 turns from Redis)
→ Answer Synthesis (Gemini 2.5 Flash, strict governance-analyst system prompt)
→ Response + Memory Write (answer + sources returned; Q&A pair saved to Redis, 30-min TTL)
Component ownership — a deliberately thin HTTP layer over a focused core:
| Component | Responsibility |
|---|---|
app.py (FastAPI) | Request validation only — delegates to the RAG core, returns JSON |
simplified_rag.py | Extraction, chunking, embeddings, Pinecone upsert/query, orchestration |
chat_engine.py | Sub-query decomposition + answer synthesis |
utils.py | Redis-backed conversation memory (TTL + window trimming) |
prompts.yaml | Every prompt template, loaded once at startup — no hardcoded prompts in code |
Key engineering decisions:
- Sub-query decomposition trades ~2–3s latency for completeness — a single embedding call averages across multi-part questions and silently drops one half of the answer; splitting into up to 5 sub-queries and retrieving in parallel fixes that. For a governance tool, correctness wins over speed.
- 1536D embeddings, chosen for the domain, not the default — policy language draws fine distinctions (prohibits vs. recommends vs. permits) that benefit from the higher-dimensional space; Pinecone's free tier absorbs the cost.
- Chunk size tuned to the corpus, not left at the library default — the default 1,200-token target collapsed each short AGORA document into a single chunk, destroying retrieval granularity. Retuning to 500 tokens (20% overlap) produced 36 meaningfully-distinct chunks from 10 files.
- Two separate Redis stores, not one — a 30-minute, 5-message window feeds the model (older context dilutes retrieval relevance); a separate 7-day store backs the UI so users can reopen old sessions. Different jobs, deliberately kept apart.
- Failures are loud, not silent — an earlier version returned zero-vectors on embedding failure, which Pinecone rejects with an opaque error. Replaced with explicit exponential-backoff retries (3 attempts) and hard failure — silent failure in an embedding pipeline is worse than a visible one.
- Namespace isolation is trust-based, not cryptographic — a single Pinecone index serves multiple document sets via
entity_idnamespaces, unvalidated server-side. Fine for this use case; called out explicitly as unsuitable for sensitive multi-tenant data without separate indexes.
Known limitations: scanned/image-based PDFs aren't extractable via pdfplumber; the 30-minute model-context TTL means very old sessions lose continuity even though UI history persists; namespace isolation isn't enforced server-side.
Next architectural steps: hybrid search (BM25 + semantic) to catch exact statute references semantic search misses; cross-encoder reranking on the top-20 before final top-4 selection; metadata filtering (jurisdiction, document_type, year) for scoped queries.
#Evaluation
A dedicated evaluation dashboard tracks retrieval and answer quality on the eval-dashboard branch, deployed live at evaluation-dashboard.streamlit.app. It gives a quantitative view of how well the system retrieves and grounds answers, rather than relying on qualitative spot-checks alone.