Web text

How do I get clean text out of a web page, without the HTML junk?

A web page is mostly markup, and the part you want to read is a small share of what the browser downloads. Fetch the page with HTML stripping switched on, so only the readable text comes back.

What you're seeing

You paste a page into a model and half of what you sent is menus, cookie banners and script. The model wades through all of it, and you pay for every word.

Why it happens

How to fix it

  1. Fetch the page with HTML stripping switched on, so only the readable text comes back.
  2. Save that text to a file you can open.
  3. If you want a summary, add a model step that reads the clean text instead of the raw page.
The recipe that does it

Turn a web page into clean text you can use

A markdown file holding the readable text of the page, with the markup stripped out.

Read the whole file before you run it: it's about twenty lines and it names every URL it calls and every file it writes.

See the recipe →

The short version

Cleaning the page before a model sees it means the model reads what a person would read. It's one setting on the fetch step.

Related questions

Does it run the page's JavaScript?

No. It reads what the server sends. Pages that build themselves in the browser need a browser step.

Is this allowed?

It reads public pages the way a browser does. Respect each site's terms and keep the frequency low.

Can it do several pages?

Yes, with one fetch step per page.

Vibe Engine (Vibra-Ingenn) runs these recipes on your own machine. Get the engine, browse the recipe library, or watch three recipes run.

Using an AI assistant? Ask your AI to check out adeptuscamini.com.

Written and maintained by Adeptus Camini, a one-person workshop. These are tools we build and run ourselves.

Other problems people hit

Chaining

How do I run several automations in order with one command?

Human checkpoint

My automation gets stuck at a login page or a CAPTCHA

Secrets

How do I stop pasting API keys into scripts and prompts?