carprices.lk
Role: Founder
Domain / Industry: Automotive / Data Analytics
Project Overview
carprices.lk tracks car listings across Sri Lanka’s biggest classifieds sites — ikman.lk and riyasewana.com — for popular models (Suzuki Alto, Toyota Hilux, Suzuki SX4, Daihatsu Terios, Toyota Rush, Toyota Vitz, Toyota Axio) and turns raw ad text into structured, comparable data. The goal is to give buyers and sellers real price trends instead of guesswork.
Technologies Used
| Category | Technologies |
|---|---|
| Backend | FastAPI, Python, SQLModel/SQLAlchemy (async), Alembic |
| AI / Extraction | Google Gemini 2.5 Flash for structured field extraction from listing text |
| Frontend | Server-rendered Jinja2 + HTMX + Tailwind CSS dashboard |
| Database | PostgreSQL (prod), SQLite (local) |
| Infrastructure | Railway (web service + cron scraper), shared Postgres |
System Architecture
A registry-driven scraper layer covers two sites (ikman.lk, riyasewana.com) behind a shared enrichment pipeline: listing pages are scraped for links, new links are fetched in parallel, and detail-page text is batched through Gemini for structured extraction (year, mileage, condition, price, etc.), overlapping fetch and LLM calls for throughput. Riyasewana’s Cloudflare protection is handled via TLS/HTTP2 browser impersonation. Results land in Postgres and are served through a lightweight HTMX dashboard for browsing listings and price trends per model.
Key Features Implemented
- Automated daily scraping of 7 popular car models across two major Sri Lankan classifieds sites
- LLM-based structured extraction from unstructured ad text (year, mileage, transmission, condition, etc.)
- Deduplication so only genuinely new listings hit the detail-page fetch and LLM pipeline
- Server-rendered dashboard with per-model tabs, filtering, and price trend views (HTMX partial swaps, no client-side JS framework)
- Cron-based scraping on Railway, with automatic schema migrations on deploy
Challenges & Solutions
- Scraping a Cloudflare-protected site by impersonating a real browser’s TLS/HTTP2 fingerprint (
curl_cffi) instead of plain HTTP requests. - Keeping LLM costs and latency down by batching listing text into fixed-size Gemini requests and pipelining extraction with remaining page fetches, rather than waiting for every fetch to finish first.
- Avoiding redundant work by checking link+price against the database before enriching or re-scraping a listing.
Learnings & Takeaways
- Practical experience combining traditional web scraping with LLM-based structured extraction at scale.
- Designing a registry pattern that lets new tracked cars/sites be added with minimal code changes.
- End-to-end ownership of a solo side project — scraping, backend, data pipeline, UI, and deployment.