Skip to content
ExtractDirect downloadVersion 0.2.0

AnyTool Smart Data Extractor

Turn visible page data into useful files.

Export HTML tables and repeated items like lists and cards from the page you are viewing to CSV, JSON or XLSX.

Download v0.2.0 (ZIP, 38.8 KB)

For Google Chrome and Microsoft Edge. Installed manually: no store, no one-click install.

Version
0.2.0
Released
October 10, 2026
File
anytool-data-extractor-0.2.0.zip
SHA-256
ff627f58867b889d26d1046654c775c3e5eb8fb59a38278fc98bcca20114b769

How to install in Chrome and Edge · How to update

The extraction workspace listing the table found on a sample page, with column checkboxes, an editable preview and CSV and JSON export buttons.

1 / 3 · The extraction workspace listing the table found on a sample page, with column checkboxes, an editable preview and CSV and JSON export buttons. Shown on a sample page.

What it is

AnyTool Smart Data Extractor turns data that is already on the page you are viewing into a file: HTML tables, and repeated structures such as product cards, result lists and link lists.

You choose what to extract, review a preview you can edit, pick the columns, and only then export to CSV or JSON. Extraction runs entirely in your browser, on the page you opened, when you ask.

It is intended for content you are authorised to access and export. It does not log in for you, solve CAPTCHAs, run in the background or crawl across sites. Optional automatic paging only starts when you press Start, stays on the one site you are on, and stops at the limit you set.

Who it is for

  • Analysts and researchers

    Move a published table or a results list into a spreadsheet without retyping, with columns you choose and a preview you can correct before saving.

  • Marketers and operations teams

    Turn a pricing grid, a directory listing or a product list on a page you can access into a CSV for a report or an import.

  • Developers and data engineers

    Prototype against real structure quickly: export repeated cards as JSON with link and image addresses, then build from it.

  • Journalists and students

    Collect figures from a public page into a file you can sort and chart, keeping the source address in the workspace header.

What it does

Features

  • One-click auto-detect of the most substantial table or list on the page
  • Finds every HTML table on the page and previews the first rows
  • Point-and-click selection of repeated elements (cards, list items, rows)
  • Text, link and image columns detected automatically, renameable and editable
  • CSV (UTF-8 with Unicode preserved), JSON and XLSX export
  • Spreadsheet formula-injection protection for CSV
  • Likely personal data columns are left unselected by default
  • Manual, page-by-page appending for paginated listings
  • Optional automatic paging through next-page links or infinite scroll, with a page limit, delay and Stop button

What it does not do

  • Start by itself, run in the background, or move to another site: automatic paging runs only after you press Start and never leaves the site you are on
  • Bypass logins, CAPTCHAs, paywalls, access controls or rate limits
  • Read form field values, passwords, cookies or storage
  • Run in the background or send extracted data anywhere

Use cases

  • Export an HTML table to CSV

    Choose Find Tables, pick the table, deselect columns you do not need, rename headers and download. Merged cells are expanded by repeating their value so every row is complete.

  • Export a list of cards to JSON

    Choose Pick Repeated Items and click one card. The extension finds the similar siblings, derives text, link and image columns, and exports clean JSON records.

  • Collect titles and URLs from a results page

    Link columns hold absolute URLs, and link text is a separate column. Use them to build a reading list or a content inventory.

  • Gather a multi-page list, manually

    Extract the first page, open the next one yourself and choose Append This Page. The same pattern is applied to the new page and the workspace grows. Nothing is loaded for you.

  • Clean up before the spreadsheet

    Edit cells in the preview, rename columns and untick noise. Edits are included in the export, so the file is ready to use.

  • Convert a specification table to JSON

    Export the same table as JSON to use as test data or configuration. Values are strings exactly as displayed.

How to use it

  1. Open the page that holds the data and click the AnyTool Smart Data Extractor toolbar icon.
  2. Choose Auto-detect Data for one click, or Find Tables, or Pick Repeated Items and click one sample item.
  3. In the workspace, review the preview, rename or deselect columns, and fix any cells.
  4. Export CSV, JSON or XLSX. For more pages, append them yourself with Append This Page, or press Start collecting to page through them automatically.

Step-by-step guide: How to export a table from a web page to CSV

How it works

  1. Tables are read structurally

    Each table becomes a rectangular grid. Rowspan and colspan are expanded deterministically by repeating the merged value, header rows are detected from thead or an all-header first row, duplicate header names are made unique, and a nested table is listed separately instead of leaking into its parent cell.

  2. Repeated items from one click

    Starting from the element you click, the extension climbs until it finds a level with at least two similar siblings, meaning the same tag and a shared class. Parent level and Child level let you correct it, and the number of items found is shown before you extract.

  3. Fields keyed by structure

    For each item, text, link addresses and image addresses are recorded against their position inside the item. Matching positions across items become columns, named from class names where possible and always renameable.

  4. Sensitive columns default to off

    Columns that look like email addresses, phone numbers, card numbers, government IDs or credentials, or whose header suggests a secret, start unselected with a note. You can include them deliberately.

  5. Exports that open correctly

    CSV follows RFC 4180 quoting with CRLF line endings, optional UTF-8 byte-order mark for Excel, and cells beginning with =, +, -, @, tab or carriage return are prefixed with a single quote unless they are plain numbers. JSON is not altered.

  6. Local and temporary

    Extraction happens in the page you opened, and the result is held in the extension's own storage only while the workspace tab is open. Form field values are never read.

Tips for better results

  • For a list that loads more as you scroll, use Collect more pages with Infinite scroll, or scroll yourself first; only what is on the page at that moment is read.
  • Start with a small page limit (3 to 5) to confirm the pattern is right before collecting more.
  • Use a delay of 2 seconds or more so you are gentle on the site you are collecting from.
  • Click a representative item that has all the fields you want, not an unusual one.
  • If the selection is too narrow or too broad, use Parent level or Child level before extracting.
  • Rename columns in the workspace to match the system you are importing into.
  • Keep formula protection on when a spreadsheet will open the CSV; turn it off only for tools that read the file programmatically.
  • Check the notes above the preview. They explain merged cells, missing headers and truncation.
  • Choose XLSX when a spreadsheet has accents or non-Latin text and you do not want to deal with CSV encoding.

Compared with other options

There are many ways to get data out of a page. They differ in what they can reach, who sees the data and how much setup they need.

OptionBest forTrade-off
Select and copy-paste into a spreadsheetA small, simple table.Merged cells, hidden text and multi-line cells often land in the wrong place; no column choice; manual for every page.
Spreadsheet import functions such as IMPORTHTMLPublic pages that update and need to stay linked.The fetch happens on the spreadsheet provider's servers, so it cannot see pages behind your login or content rendered by scripts.
Online scraping servicesLarge, scheduled collection from public sites at scale.Data and URLs pass through a third party; they may conflict with a site's terms and cannot use your session.
Writing a script in the DevTools consoleDevelopers who want full control.Takes time to write and maintain; no preview or column UI.
AnyTool Smart Data ExtractorOne-off or occasional exports from pages you are already viewing, with preview, column choice and safe CSV.Paging is opt-in and bounded (up to 50 pages per run); no scheduling, cross-site crawling or background extraction; 20,000-row cap.

Permissions explained

The extension asks for the following browser permissions and nothing else. It requests no access to all websites.

PermissionWhy it is needed
activeTabLets the extension read the page you are viewing, only after you click the toolbar icon.
scriptingRuns the table finder and the repeated-item picker in the current tab.

Optional site access (on request only). Only if you press Start collecting for automatic paging: your browser asks for access to the one site you are on, and it is released when the run ends. It is never requested at install and never covers all sites.

Privacy

What it reads

  • The text, link addresses and image addresses of the tables or repeated items you choose to extract.
  • Never the values of form fields, passwords, cookies, local storage or hidden inputs.
  • If you start automatic paging: the same fields from further pages of the same site, which your browser loads in your own tab at your request.

What it keeps

  • Extracted records are held in the extension's own IndexedDB storage while the workspace tab is open, and deleted when it closes. Anything abandoned is removed after six hours.
  • Column names and selections are kept in memory only.

Nothing it reads is sent to AnyTool Studio or any other server. Read the full privacy notice.

Browser compatibility

  • Google Chrome

    Manifest V3, Chrome 116 or later.

    Install from the ZIP download above. Not currently listed in the browser store.

  • Microsoft Edge

    Manifest V3, Edge 116 or later.

    Install from the ZIP download above. Not currently listed in the browser store.

Limitations

  • Automatic paging needs a next-page link or an infinite-scroll list that matches the pattern you picked. It stops at your page limit (up to 50 pages per run), at 20,000 rows, or when nothing new appears.
  • Tables built from divs are not tables; use Pick Repeated Items for those.
  • Content inside cross-origin iframes is not reachable.
  • Merged cells (rowspan/colspan) are expanded by repeating the cell's value in each covered position.
  • Very large pages are capped (20,000 rows per extraction) with a visible warning.

Extension or online tool?

Extension. Use the extension to pull a table or a list of cards from a page you are viewing, including pages that need you to be signed in.

Online tool. Use an online tool when you already have a CSV or JSON file and want to convert, merge or clean it.

Troubleshooting

No tables were found.

The data may not be in a table element. Use Pick Repeated Items and click one row or card.

The repeated-item picker selected the wrong group.

Use the Parent and Child buttons in the picker bar to move the selection up or down a level, or click a different sample.

The page loads more rows as I scroll.

Scroll to load the rows you want first, then extract. The extension only reads what is on the page at that moment.

Frequently asked questions

Is extracted data sent to a server?

No. Extraction and export happen in your browser, and the extension makes no network requests.

Will the CSV open correctly in Excel with accents and non-Latin text?

Yes. CSV files are UTF-8 with a byte-order mark by default, which Excel uses to detect the encoding. You can turn the mark off.

What is formula-injection protection?

Cells that begin with =, +, -, @, tab or carriage return are prefixed with a single quote so a spreadsheet shows them as text instead of running them. Plain negative numbers are left alone. JSON export is not altered.

Can it scrape a whole site?

No. It works on the site in front of you, when you ask. There is no scheduler, background extraction or crawling across sites. Optional automatic paging only moves through that site's result pages after you press Start, up to the limit you set.

Why are some columns unchecked?

Columns that look like email addresses, phone numbers, card numbers or secrets start unselected so personal data isn't exported by accident. You can select them deliberately.

Can it follow pagination automatically?

Yes, if you choose to. In the workspace, press Start collecting, set a page limit and a delay, and pick next-page links or infinite scroll. Your browser asks for access to that one site for the duration of the run and it is released afterwards. It never starts by itself, stays on the same site, and you can press Stop at any time. You can also append pages manually.

Does it support infinite scroll?

Yes, as an opt-in mode. Choose Infinite scroll under Collect more pages; it scrolls your tab to the end, waits for the delay you set, reads the items again and stops when nothing new appears or at your limit.

Why does my browser ask for access to a site?

Only when you press Start collecting. Moving through pages needs access to the one site you are on, so the extension asks for exactly that site and gives it back when the run ends. It never asks at install and never asks for all sites.

How does auto-detect choose?

It scores visible tables and groups of similar items by how much data they hold and how many fields each item has, ignoring navigation and footers, and opens the highest scorer. If it is wrong, use Find Tables or Pick Repeated Items.

Can formulas run in the XLSX export?

No. Every cell is stored as text, so a value such as =HYPERLINK(...) is shown as text rather than evaluated.

What happens to merged cells?

The merged value is repeated in every position it covered, and the workspace tells you the table had merged cells.

Is the data typed?

Everything is exported as text exactly as displayed, in both CSV and JSON. Convert types in your own tooling.

How many rows can it handle?

Up to 20,000 rows per extraction, with a visible warning if a page has more.

Why is an email or phone column unchecked?

To avoid exporting personal data by accident. Tick it if you are allowed to export it.

Does it work on pages behind a login?

It reads what your browser displays, so yes. You are responsible for being allowed to export that content; it does not bypass access controls, CAPTCHAs or rate limits.

Is anything stored after I close the workspace?

No. The extracted data is deleted when the workspace tab closes, and anything abandoned is cleared after six hours.

Which browsers does it support?

Google Chrome and Microsoft Edge, version 116 or later. It was tested in Chromium and Google Chrome; an Edge-specific test pass is still to be done.

  • CaptureDirect download

    AnyTool Screenshot

    Capture more than what's on your screen.

    Capture the visible area, a region or a whole scrolling page; save PNG, JPEG or PDF, or copy. Processed on your device.

    • Visible-area capture
    • Drag-to-select region capture
    • Full scrollable page capture with progress and cancel, including scrolling panels and same-origin iframes

    View details

  • InspectDirect download

    AnyTool Font & CSS Inspector

    Understand the design behind any element.

    Click any element for fonts, colors, contrast, spacing and CSS variables, plus a page palette. Copy values or CSS.

    • Hover highlight with a size and tag tooltip
    • Click to select, arrow keys to move to parent or child
    • Font family, size, weight, style, line height and letter spacing

    View details