Files
Shubhamsaboo c2273fff20 docs: rewrite web_scraping_ai_agent README to match shipped code
The folder only ships a local ScrapeGraphAI implementation (ai_scrapper.py,
local_ai_scrapper.py), but the README documented an entire "Cloud SDK"
version referencing a scrapegraph_ai_sdk/ folder and quickstart.py /
smart_scraper_demo.py / scrapegraph_app.py — none of which exist anywhere
in the repo.

Remove all of that imaginary content and reframe as the local-only agent
it actually is. Also fix the clone-path typo (web_scrapping -> web_scraping)
and drop the now-inaccurate "(Local & Cloud SDK)" label in the root README.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-12 19:57:42 -07:00

2.3 KiB

🕷️ Web Scraping AI Agent

🎓 FREE Step-by-Step Tutorial

👉 Click here to follow our complete step-by-step tutorial and learn how to build this from scratch with detailed code walkthroughs, explanations, and best practices.

AI-powered web scraping using ScrapeGraphAI - extract structured data from websites using natural language prompts. This agent runs locally with the open-source scrapegraphai library.


📁 What's Inside

Files: ai_scrapper.py, local_ai_scrapper.py

Use the open-source ScrapeGraphAI library that runs on your local machine.

Pros:

  • Free to use (no API costs)
  • Full control over execution
  • Privacy-friendly (all data stays local)

Cons:

  • Requires local installation and dependencies
  • Limited by your hardware
  • Need to manage updates

🚀 Getting Started

  1. Clone the repository
git clone https://github.com/Shubhamsaboo/awesome-llm-apps.git
cd awesome-llm-apps/starter_ai_agents/web_scraping_ai_agent
  1. Install dependencies
pip install -r requirements.txt
  1. Get your OpenAI API Key
  1. Run the Streamlit App
streamlit run ai_scrapper.py
# Or for local models:
streamlit run local_ai_scrapper.py

💡 Use Cases

E-commerce Scraping

# Extract product information
prompt = "Extract product names, prices, and availability"

Content Aggregation

# Convert articles to structured data
prompt = "Extract article title, author, date, and main content"

Competitive Intelligence

# Monitor competitor websites
prompt = "Extract pricing, features, and updates"

Lead Generation

# Extract contact information
prompt = "Find company names, emails, and phone numbers"

🔧 How It Works

  1. You provide your OpenAI API key
  2. Select the model (GPT-4o, GPT-5, or local models)
  3. Enter the URL and extraction prompt
  4. The app uses ScrapeGraphAI to scrape and extract data locally
  5. Results are displayed in the app

📖 Documentation