Tutorials
Step-by-step tutorials on Python scrapers, Reddit APIs, Google Trends, RAG pipelines, and job data. Real production code, tested patterns.
81 articles
How to Scrape CutShort Jobs for India Tech Hiring Data (No API)
CutShort is where Indian startups post engineering and product roles, and it has no public jobs API. Learn how to extract titles, companies, salary ranges, skills, and experience bands as structured JSON for recruiting feeds and talent market research.
Airbnb Scraper: Listing Prices, Ratings, and Availability Without the API
Airbnb has no public listings API. Here's how to pull nightly price, rating, superhost status and coordinates from Airbnb search results using Python.
Maps Leads: Verified Email Extraction from Google Maps Business Listings
How to pull B2B leads from Google Maps with MX-verified emails, past the official API's 120-result cap, and pay only for contactable businesses.
Realtor.com Scraper: Property and Agent Data Without the MLS Paywall
Scrape Realtor.com listings in Python — price, beds, baths, county, and the listing agent + brokerage office on every record. No login, no API key.
Redfin Scraper: For-Sale and Sold Property Data in Python (No API Key)
Scrape Redfin listings by city, ZIP, or URL in Python — price, beds, baths, sqft, agent, MLS status, and coordinates. No login, no browser, no unblocker.
Trustpilot Business Search: Discover Companies by Category with TrustScore Data
Search Trustpilot by category or keyword in Python to build company lists with TrustScore, review count and verification status — no login required.
Zillow Property Details API: Zestimate, Tax History, and Agent Data by URL
Pull full Zillow property records by URL or ZPID in Python — price history, tax history, schools, HOA, and agent data the official API never exposed.
Zillow Rental Listings API: Rent Estimates and Availability at Scale
How to pull Zillow for-rent listings by location in Python — monthly rent, availability date, beds, baths, and Rent Zestimate — with no login or API key.
How to Scrape Zillow Listings in Python (2026 Guide, No Blocking)
Scrape Zillow for-sale listings by city or ZIP in Python — price, beds, baths, Zestimate, and days on market. No API key, no browser, no blocking.
Amazon Product Research in Python: ASIN Data, BSR and Price Without the SP-API
How to scrape Amazon product listings by ASIN or keyword — price, Best Sellers Rank, star rating, review count, seller and feature bullets — across 8 marketplaces.
Bulk Email Verification in Python: MX, SMTP and Catch-All Detection Without an API
How to verify thousands of email addresses for free using DNS MX lookups and optional SMTP handshakes. No API key, no monthly subscription.
Scrape Business Emails, Phones and Social Profiles from Any Website
How to extract contact information from a list of company domains at scale — emails, phone numbers, LinkedIn, X, Instagram and more — with no API key.
TripAdvisor Reviews Scraper: Hotels, Restaurants and Attractions Without an API
How to pull TripAdvisor reviews at scale — star rating, full text, trip type, owner response and sub-ratings for any listing — with no TripAdvisor API key.
YouTube Channel Scraper: Subscribers, Video Stats and Channel Data Without a YouTube API Key
How to pull YouTube channel metadata and video lists — subscriber count, view counts, publish dates, video descriptions — without using the YouTube Data API or paying for quota.
Company KYB Resolver: Resolve LEI, EU VAT, and SEC CIK in One API Call
KYB in one call. Resolve a company name to its GLEIF LEI, EU VAT status, and SEC EDGAR CIK. No API key required. Built for onboarding and counterparty verification.
EU VAT Validator: Bulk VIES Verification Without an API Key
Validate EU VAT numbers in bulk via the official VIES service. Returns validity, registered company name, and address. No API key required. Zero charge on empty runs.
GLEIF LEI Lookup API: Resolve Legal Entity Identifiers Without an API Key
Look up Legal Entity Identifiers (LEIs) and registered company data from the GLEIF global registry. Search by company name or LEI code. No API key required.
The Mine Works MCP Server: Add 29 Data Tools to Claude Desktop in 2 Minutes
Connect LinkedIn, Reddit, SEC filings, B2B leads, PubMed, arXiv, Google Trends, and more to Claude Desktop or Cursor via a single MCP endpoint. No subscriptions — billing flows through your Apify account.
arXiv Scraper: Search AI, Physics, and Biology Preprints with PDF Links via API
Search arXiv preprints by keyword, category, or author. Returns title, abstract, authors, categories, and PDF links. Track AI research before peer review. No API key required.
B2B Leads Finder: Business Emails and LinkedIn Profiles Without Apollo or ZoomInfo
Find business emails, LinkedIn profiles, and job titles for decision-makers at target companies. No Apollo, ZoomInfo, or Lusha API key required. $0.003 per lead.
ClinicalTrials Bulk Exporter: 575K Trials Filtered by Condition, Phase, and Status
Download structured records from ClinicalTrials.gov filtered by disease condition, trial phase, enrollment status, intervention type, or country. No API key required.
ClinicalTrials Sponsor Intelligence: Pharma Pipeline Tracking with FDA Approval Cross-Reference
Track clinical trials by sponsor name or condition, with optional cross-referencing against FDA drug approvals to identify which interventions received regulatory clearance.
CMS Hospital Quality Data: Star Ratings, HCAHPS Scores, and Complication Rates via API
Pull CMS hospital quality metrics for 4,500+ US hospitals: overall star ratings, patient satisfaction scores, complication rates, and readmission rates. No API key required.
Company KYB Resolver: LEI, EU VAT, and SEC CIK Lookup in One API Call
Resolve a company name to its Legal Entity Identifier (LEI), EU VAT registration status, and SEC EDGAR CIK in a single call. No API key required. Built for KYB onboarding and counterparty verification.
FDA 510(k) Clearances API: Medical Device Intelligence by Company, Device, or Product Code
Search FDA 510(k) premarket clearances by company, device name, or product code. Returns applicant, device, decision, and dates for medtech competitive intelligence. No API key required.
LinkedIn Company Scraper: Size, Industry, Website, and Followers Without Login
Scrape LinkedIn company pages for employee count, industry, HQ, founding year, website, follower count, and specialties. No LinkedIn login or API key required.
LinkedIn Jobs Scraper: Job Listings by Keyword and Location Without Login
Scrape LinkedIn job listings by keyword and location: job title, company, seniority, applicant count, and job description. No LinkedIn login required.
LinkedIn Post Search: Find Posts by Keyword Without Login
Search LinkedIn posts by keyword and return author profiles, headlines, and post snippets. Uses Google indexing as the data surface — no LinkedIn login or cookies needed.
LinkedIn Profile Scraper: Experience, Education, and Skills Without Login
Extract full LinkedIn profile data — work history, education, skills, connections — from a list of profile URLs. No LinkedIn cookies or login required.
Medicare Part D Drug Spending Data: Unit Cost, Claims, and Beneficiary Counts via API
Pull Medicare Part D drug spending from CMS: total cost, claims, beneficiaries, and unit price by drug name or manufacturer. Track pharmaceutical pricing trends. No API key required.
NIH RePORTER API: Search Grant Funding, Award Amounts, and Principal Investigators
Search NIH grant awards by topic, agency, institution, or state. Returns project title, abstract, award amount, and PI data for research funding intelligence. No API key required.
OpenCitations API: Build Citation Graphs from 1.6 Billion Open Citation Links
Pull all citing papers and cited references for any DOI from OpenCitations' 1.6B citation index. Build citation networks, map research influence, and run bibliometric analysis without a subscription.
openFDA Scraper: Drug Adverse Events, Device Recalls, and Food Safety Data
Extract FDA drug adverse events, device recalls, 510k clearances, and food enforcement actions from all openFDA endpoints using a single Python script. No API key required.
PubMed Scraper: Search 36M Biomedical Articles, Abstracts, and MeSH Terms via API
Search PubMed for biomedical literature by query, author, or MeSH term. Returns PMID, title, abstract, authors, journal, and DOI. No API key required. Ideal for systematic reviews and RAG pipelines.
AliExpress Product Data API: Prices, Ratings, and Orders in Python
AliExpress affiliate API has restricted coverage. Learn how to scrape AliExpress product listings for prices, ratings, order counts, and seller data as structured JSON — no affiliate approval needed.
How to Scrape AmbitionBox Company Reviews and Ratings
AmbitionBox is India largest employer review platform with 300,000 companies. Learn how to pull ratings, review counts, salary data, and dimension scores as structured JSON without any official API.
ClinicalTrials.gov API v2: How to Search 500,000 Studies and Track Trial Status
ClinicalTrials.gov upgraded to a v2 REST API in 2024. Here is how to use it, what changed from v1, and how to build automated trial monitoring pipelines in Python.
CourtListener API: How to Search US Court Records and Case Law Programmatically
CourtListener exposes 10M+ court opinions and dockets via a free REST API. Here is how to query it, what the rate limits actually are, and when a scraper is faster.
Crossref API: 150 Million DOIs, Citation Counts, and Bibliographic Data for Free
Crossref is the canonical DOI resolver for 150M+ scholarly works. The REST API returns publication metadata, reference lists, and citation counts with no authentication.
How to Scrape Crunchbase Company Profiles in Python (Funding, Investors, No API Key)
Crunchbase has no free public API. Learn how to extract company names, total funding raised, investor counts, latest rounds, and headquarters as structured JSON using an Apify scraper with residential proxy bypass.
FDA Recall Data API: How to Monitor Drug, Device, and Food Recalls Programmatically
openFDA exposes drug recalls, device recalls, and food safety enforcement actions via a REST API. Here is how the endpoints work and what the data actually contains.
Federal Register API: How to Track US Rules, Proposed Rules, and Executive Orders
The Federal Register publishes every US executive action, proposed rule, and final rule via a REST API. Here is how to query it and what the data contains.
How to Scrape Google News in Python (No API Key Required)
Google killed its News API in 2013. Learn how to pull headlines, sources, and publication dates from Google News in Python using the RSS feed, the GNews approach, and a pay-per-result scraper.
Google Trends API for Python in 2025: pytrends vs Scraper
Google Trends has no official API. Learn why pytrends breaks, how the SERP API approach works, and the fastest way to pull trend data into Python without getting rate-limited.
India Government Data API: How to Pull Any data.gov.in Dataset Without the Documentation Confusion
data.gov.in has 10,000+ datasets including mandi prices, foreign trade, and census data. The OGD API works but has quirks that are not documented anywhere.
How to Scrape IndiaMART B2B Suppliers in Python (Phone, Price, Leads)
IndiaMART is India's largest B2B marketplace with no public API. Learn how to extract supplier names, phone numbers, cities, product categories, and price indications as structured JSON for sales prospecting and market research.
Instagram Profile Data Without the Meta API: Followers, Bio, and Posts at Scale
Meta restricts the Instagram Graph API to your own accounts. For researching public third-party profiles at scale, here is what data is available and how to collect it.
How to Scrape JustDial Business Listings in Python (Phone, Address, Ratings)
JustDial is India's largest local business directory with no public API. Learn how to extract business names, phone numbers, addresses, geo-coordinates, and review data by city and category as structured JSON.
How to Scrape LinkedIn Employees Without Login or Sales Navigator
LinkedIn has no public API for employee data. Learn how to pull B2B leads, employee lists, and org chart data from LinkedIn company pages without a LinkedIn account or Sales Navigator subscription.
How to Scrape the Meta Ad Library in Python (Facebook and Instagram Ads, No Login)
The Meta Ad Library has no official scraping API. Learn how to extract ad copy, creative URLs, advertiser details, platforms, and run dates from Facebook and Instagram ads by keyword or advertiser, as structured JSON.
How to Scrape Naukri.com Jobs in Python (Structured JSON with Salaries)
Naukri.com has no public API. Learn how to scrape India's #1 job board for titles, companies, salary ranges, skills, experience, and work mode as structured JSON with pay-per-result pricing.
How to Query Norway BRREG Business Register in Python (Companies, Officers, AML)
Norway's BRREG business register is public and free, but navigating the API to get companies plus officer roles requires multiple calls per entity. Learn how to extract company status, industry, address, and full officer rosters as structured JSON.
OpenAlex API: 250 Million Research Papers, Free, No Rate-Limit Workarounds Needed
OpenAlex replaced the defunct Microsoft Academic Graph with 250M+ scholarly works. The API is free, well-documented, and returns structured data including citations and author affiliations.
NPI Registry API: How to Look Up Any US Healthcare Provider Programmatically
CMS publishes the National Provider Identifier registry as a free API. Here is how to search by provider name, specialty, location, and NPI number — and what the data contains.
How to Scrape Pinterest Profiles in Python (Followers, Pins, Boards Without Login)
Pinterest has no public API for scraping profiles. Learn how to extract follower counts, monthly views, board lists, recent pins with save counts, and profile metadata as structured JSON without login or API key.
SEC EDGAR Full-Text Search API: Search Every Filing Since 2001
The free EFTS endpoint searches every SEC filing since 2001. The parts that break scripts: a required User-Agent, 100 hits per page, and a hard 10,000-result cap.
Socrata API: How to Pull CDC, HHS, NYC, and 200+ Government Data Portals
Socrata powers data portals for the CDC, HHS, Chicago, New York City, Texas, and 200+ other government entities. One API, same query syntax, all of them.
Threads Has No Public API in 2026 — Here Is How to Get Post and Profile Data Anyway
Every field you can collect from Threads without an API: post text, engagement counts, media, profiles. No login required, from $1 per 1,000 posts — plus a comparison of the scrapers that do it.
How to Scrape Trustpilot Reviews by Company Domain (Python Guide)
Trustpilot has no public API for review data. Learn how to pull business reviews, star ratings, trust scores, and business replies from any Trustpilot company page using Python.
USASpending.gov API: How to Pull Federal Contracts, Grants, and Awards Programmatically
USASpending.gov tracks every federal dollar spent. The API is public and free but the endpoint structure is non-obvious. Here is how to actually use it in Python.
World Bank API in Python 2025: GDP, Inflation, and 1,400 Indicators Without the SOAP Hell
The World Bank has a REST API but it returns XML by default, uses quirky pagination, and has undocumented quirks. Here is how to actually use it in Python.
World Bank Trade Data API: How to Pull Global Import and Export Statistics
The World Bank WITS database covers bilateral trade flows between 200+ countries. Here is how to access it programmatically and what the data actually contains.
How to Scrape Yellow Pages US Business Listings in Python (Phone, Address, Website)
YellowPages.com has no public API. Learn how to extract US local business names, phone numbers, addresses, websites, and categories by city and business type as structured JSON for sales prospecting and local research.
Pull SEC Filings into a RAG Pipeline with Claude and the SEC EDGAR Scraper
How to turn 10-K, 10-Q and 8-K filings into a clean, chunked, citation-grounded knowledge base an LLM can answer questions over
Scraping Reddit Comments and Full Thread Trees in 2025
Reddit's nested comment structure is complex to collect correctly. This guide covers the complete API approach for deep comment trees, deleted comments
How to Export Google Trends Data at Scale for Market Research
Exporting Google Trends for dozens or hundreds of keywords while avoiding rate limits, handling the normalization quirks
The Agentic Data Stack 2025: How to Pick the Right Scrapers for Your AI Workflow
A practical guide to building grounded AI agents with real-time scraped data. Which data sources matter for which agent types
Building a RAG Pipeline on SEC EDGAR Filings: A Step-by-Step Guide
How to scrape SEC EDGAR filings, chunk them for vector search, and build a provenance-aware Q&A system that cites specific filing sections using Claude.
Building an Automated Naukri Job Alert System with Python
How to build a custom Naukri job monitoring system that filters by salary, location, and skills — and sends instant alerts when relevant jobs post.
Build a Social Listening Agent for Threads with Claude
Use Apify's Threads Scraper with Claude to automate trend detection, brand monitoring, and content ideation from Meta's Threads platform.
Build a Custom Knowledge Base Chatbot with Claude and the RAG Crawler
Use Apify's RAG Crawler to ingest any website into a vector database, then wire Claude to answer questions against it.
Build an India Job Market Intelligence Tool with Claude and the Naukri Scraper
Use Apify's Naukri Jobs scraper with Claude to automate salary benchmarking, skills demand analysis, and hiring trend tracking for the Indian tech market.
Build a Talent Intelligence System with Claude and ATS Job Scrapers
Combine Greenhouse, Lever, and Ashby job data with Claude to automate candidate sourcing research, salary benchmarking, skills gap analysis
Automate SEO Research and Content Strategy with Claude and Google Trends Pro
Use Apify's Google Trends Pro actor with Claude to build an autonomous content calendar generator, keyword opportunity finder
Build a Reddit Intelligence Agent with Claude and the Reddit Scraper
How to combine Apify's Reddit Scraper with Claude to build an autonomous brand monitoring agent, sentiment analysis pipeline
How to Aggregate Job Postings from 500+ Companies Using Public ATS APIs
Greenhouse, Lever, and Ashby expose zero-auth public job board APIs. This guide shows how to build a job aggregator that pulls from all three and
How to Build a RAG Pipeline Using Web-Scraped Content
A complete guide to turning any website into LLM context — from crawling and chunking to embedding, retrieval, and keeping the index fresh.
How to Scrape Meta Threads Data in 2025 (Without Getting Blocked)
Meta Threads has no public API for third-party developers. This guide shows the current working approaches for extracting profile data, post content
Naukri API 2025: How to Programmatically Access India's Largest Job Board
Naukri has no public API. This guide covers the session-warming approach that bypasses Akamai bot detection
Google Trends API Python 2025: Why pytrends Keeps Breaking (and What to Use Instead)
pytrends has been unreliable for years. We explain why Google Trends blocks HTTP clients, and show you three approaches that actually work in 2025.
How to Scrape Reddit Without an API Key in 2026
The old reddit.com .json endpoints now return 403 and commercial API access is enterprise-only. Every method that still works in 2026 — with code you can use today.