Intelligent Q–A.
Answers with somewhere to point. A retrieval pipeline that transforms a website into a searchable, source-grounded knowledge base.
- FastAPI
- RAG
- BM25
- TF-IDF
- Python
The problem
behind the project.
A useful answer should remain connected to its source. This system crawls a target website, builds an index, and answers questions using material from that indexed domain.
How it works.
- 01
Crawl and prepare
A domain-bounded crawler collects website content. The pipeline chunks that material and prepares it for indexing.
- 02
Retrieve relevant content
BM25 and TF-IDF retrieval select relevant indexed content for a question.
- 03
Answer with sources
The RAG workflow produces source-attributed answers restricted to the indexed domain, exposed through a FastAPI REST API.
Engineering choices.
Keep the source boundary explicit
Restricting the workflow to an indexed domain gives the retrieval stage a clear scope and keeps source attribution central to the answer.
Separate the pipeline stages
Dedicated /crawl, /index, /ask, and /crawl-and-ask endpoints support individual stages as well as the end-to-end workflow.
Build in observability
Recall@k tests, timing metrics, and structured logging provide evidence for iterative retrieval and latency tuning. No benchmark results are claimed here.
Explore the source.
The repository is the implementation reference. The illustrations on this page explain the architecture; they are not application screenshots or measured results.
- FastAPI endpoints for crawling, indexing, and questions
- Domain-restricted BM25 and TF-IDF retrieval
- Recall@k tests, timing instrumentation, and structured logs