Google Maps Scraper screenshot thumbnail

Google Maps Scraper

The leading open-source tool for extracting business leads from Google Maps at production scale. The Go-based engine processes approximately 120 places per minute with optimized concurrency, extracting 33+ data points per listing including business name, address, phone number, website URL, rating, review count, latitude and longitude, opening hours, price level, and optionally crawling business websites for email addresses. Three interfaces serve different workflows: the CLI accepts query files for cron jobs and CI/CD pipelines with output to CSV, JSON, PostgreSQL, S3, or LeadsDB; the Web UI provides a browser-based dashboard with real-time job monitoring, a map view of scraped places, and interactive query submission; and the REST API at /api/v1 enables programmatic integration with full Swagger documentation at /api/docs. Built-in proxy rotation supports SOCKS5, HTTP, and HTTPS with authentication for large-scale runs, while the architecture scales from a laptop to Kubernetes clusters with queue-based worker distribution. The SaaS edition adds multi-user access with API key management, admin UI with 2FA, job queue orchestration, and one-command cloud deployment via an interactive wizard. An AI Agent Skill enables coding agents to run scrapes programmatically. Deploy via Docker or build from source requiring Go 1.26.5+. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.

Deploy
Steel Browser screenshot thumbnail

Steel Browser

With over 7,400 GitHub stars and benchmarked at 0.89 seconds average session lifecycle — 1.7x to 9x faster than competing browser automation platforms — Steel Browser delivers production-grade headless Chrome infrastructure purpose-built for AI agents that need to interact with the modern web. The TypeScript-based server exposes a REST API providing on-demand browser sessions with full CDP (Chrome DevTools Protocol) access, allowing connections from Puppeteer, Playwright, or Selenium through standard WebSocket endpoints without framework lock-in. Each session maintains persistent state including cookies, localStorage, IndexedDB, and authentication credentials across requests, enabling stateful multi-step agent workflows that survive session restarts. Built-in anti-detection includes stealth plugins, browser fingerprint randomization, and configurable user-agent rotation, while the proxy chain manager handles IP rotation through residential, datacenter, or custom proxy pools. CAPTCHA solving integrates natively so agents encounter fewer blocking interrupts during autonomous navigation. The Session Viewer provides real-time WebRTC-streamed visual debugging of live sessions and playback of recorded sessions with full network request logging. Browser Tools APIs convert any page to clean Markdown, readability-optimized text, PDF documents, or high-resolution screenshots with a single API call. The MCP Server integration exposes Steel sessions as tools accessible to Claude, Cursor, and other Model Context Protocol-compatible AI agents. Deploy via Docker with a single container or use Docker Compose for production configurations with automatic resource cleanup and session lifecycle management. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.

Deploy
HeadlessX screenshot thumbnail

HeadlessX

With 2,000 GitHub stars and 10 releases since its September 2025 launch, HeadlessX delivers a self-hosted browser automation platform that replaces Chromium-based scraping with Camoufox — a Firefox fork performing kernel-level fingerprint spoofing to achieve 0% detection across Cloudflare, DataDome, PerimeterX, and other anti-bot systems where Puppeteer and Playwright regularly fail. The web dashboard provides workspace-based job organization with a visual interface for configuring scrape targets, managing browser profiles, monitoring queue status, and viewing extracted results in real time. The protected REST API accepts requests with API key authentication for programmatic access, supporting HTML extraction, screenshot capture, PDF generation, and structured data parsing with configurable stealth parameters. Profile-based scraping maintains persistent browser contexts with cookie jars, localStorage, and fingerprint configurations that survive between requests — reducing cold-start latency from 25 seconds to under 2 seconds on subsequent requests. Queue-backed workflows enable batch processing of URLs with configurable concurrency, retry logic, and webhook notifications on completion. The Google AI Search integration provides AI-assisted web research workflows through dedicated endpoints. Remote MCP support exposes automation capabilities as tool endpoints for AI agent integration. Deploy via the official CLI with `headlessx init` and `headlessx start` commands, scaffolding a Docker Compose stack with Caddy reverse proxy for automatic HTTPS. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. MIT licensed.

Deploy