← Back to selected work
Selected project / 02Information retrieval / Applied AI

Intelligent Q–A.

Answers with somewhere to point. A retrieval pipeline that transforms a website into a searchable, source-grounded knowledge base.

  • FastAPI
  • RAG
  • BM25
  • TF-IDF
  • Python
Conceptual architecture illustrationDesigned to explain the system
01 / PURPOSE

The problem
behind the project.

A useful answer should remain connected to its source. This system crawls a target website, builds an index, and answers questions using material from that indexed domain.

02 / ARCHITECTURE

How it works.

  1. 01

    Crawl and prepare

    A domain-bounded crawler collects website content. The pipeline chunks that material and prepares it for indexing.

  2. 02

    Retrieve relevant content

    BM25 and TF-IDF retrieval select relevant indexed content for a question.

  3. 03

    Answer with sources

    The RAG workflow produces source-attributed answers restricted to the indexed domain, exposed through a FastAPI REST API.

03 / DECISIONS

Engineering choices.

Keep the source boundary explicit

Restricting the workflow to an indexed domain gives the retrieval stage a clear scope and keeps source attribution central to the answer.

Separate the pipeline stages

Dedicated /crawl, /index, /ask, and /crawl-and-ask endpoints support individual stages as well as the end-to-end workflow.

Build in observability

Recall@k tests, timing metrics, and structured logging provide evidence for iterative retrieval and latency tuning. No benchmark results are claimed here.

04 / IMPLEMENTATION

Explore the source.

The repository is the implementation reference. The illustrations on this page explain the architecture; they are not application screenshots or measured results.

  • FastAPI endpoints for crawling, indexing, and questions
  • Domain-restricted BM25 and TF-IDF retrieval
  • Recall@k tests, timing instrumentation, and structured logs
View GitHub repository
Next project FSapp.