Research-ready extraction

Turn case studies into structured datasets.

Provide a few example pages and the fields you need. Case Study Scraper finds pages with matching URL patterns and extracts the information into CSV or JSON. It is designed for case studies, customer stories, and recent-project pages.

Run configuration

Input

Target Site

northstar-studio.example

Example Pages

  • /work/luma-coffee
  • /work/fieldwork-health
  • /work/orbit-logistics

Extraction Fields

  • Client
  • Industry
  • Services
  • Outcome

Extraction Dataset

2 of 3 Extraction Results shown

CSVJSON
  1. RECORD 01Successful
    Client
    Luma Coffee
    Industry
    Hospitality
    Services
    Brand strategy, web design
    Outcome
    A clearer digital launch

    Canonical Page URL: https://northstar-studio.example/work/luma-coffee

  2. RECORD 02Successful
    Client
    Fieldwork Health
    Industry
    Healthcare
    Services
    Research, product design
    Outcome
    A simpler patient journey

    Canonical Page URL: https://northstar-studio.example/work/fieldwork-health

A fictional Northstar Studio run configuration with a Target Site, three Example Pages, and four Extraction Fields becomes source-linked Extraction Results, with two example results shown, available as CSV or JSON.

How it works

From a few examples to a reviewable dataset.

Configure the page pattern and fields once. AI-assisted matching and extraction help turn repeated project pages into consistent Extraction Results.

  1. 01

    Show what matches

    Provide the Target Site and a few Example Pages that demonstrate the relevant URL structure.

  2. 02

    Choose what to collect

    Define the Extraction Fields you want from each Matching Page, such as client, industry, services, and outcome.

  3. 03

    Review and download

    Inspect successful or failed Scrape Jobs, then download the successful Extraction Results as CSV or JSON.

Output

A sourced Extraction Dataset, ready for analysis.

Each successful Extraction Result keeps the structured values you chose alongside its Canonical Page URL, so the result remains easy to trace and review.

Successful Extraction Results
Structured values under your chosen fields
Source linked
  • Luma Coffee

    Hospitality · Brand strategy, web design · A clearer digital launch

    https://northstar-studio.example/work/luma-coffee

  • Fieldwork Health

    Healthcare · Research, product design · A simpler patient journey

    https://northstar-studio.example/work/fieldwork-health

CSV

Columns use your user-facing Field Labels for familiar spreadsheet analysis.

JSON

Properties use stable Field Keys for reliable downstream workflows.

Review outcomes before download

Inspect successful and failed Scrape Job outcomes first. Downloads become available after the run is terminal and contain successful Extraction Results.

SuccessfulFailed

Use cases

Built for case studies. Flexible enough for other repeated page types.

Apply the same example-led workflow when a site publishes structured information across a consistent family of pages.

01

Customer stories

client, industry, services, outcomes.

02

Project portfolios

client, location, project type, scope.

03

Team profiles

name, role, specialties.

04

Location pages

address, region, contact details.

Build the dataset you need

Turn repeated project pages into structured, source-linked data.

Start with a few examples, choose your fields, and review the results before downloading CSV or JSON.