# Smart Extraction for scraping

**URL:** <https://forum.gumloop.com/t/smart-extraction-for-scraping/1562>\
**Category:** Feature Request\
**Tags:** Extract-Data, Website-Scraper\
**Created:** [March 1, 2025, 8:26am UTC](https://forum.gumloop.com/t/smart-extraction-for-scraping/1562 "2025-03-01T08:26:08Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Shrikar](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.gumloop.com/shrikar/32/300_2.png) [@Shrikar](https://forum.gumloop.com/u/Shrikar)\
**Post date:** [March 1, 2025, 8:26am UTC](https://forum.gumloop.com/t/smart-extraction-for-scraping/1562/1 "2025-03-01T08:26:08Z")

</div>

What would be really cool is this workflow

- Fetch Raw html content
- Let end user define the scope of the html element and where the data exists
- Run llm on that short dom element  
Looks like we are running the llm on the whole page content that is probably waste of token and hence the high usage of credits for web agent scraper?

More on this here: [Web scraping node - Return HTML - #7 by Shrikar](https://forum.gumloop.com/t/web-scraping-node-return-html/1554/7)

---

<div class="post-metadata">

**Author:** ![system](https://canada1.discourse-cdn.com/flex027/uploads/gumloop/original/1X/d02ff8efcde6ac5d3ac8aa0376137d5f804da549.png) [@system](https://forum.gumloop.com/u/system)\
**Post date:** [March 5, 2025, 8:26pm UTC](https://forum.gumloop.com/t/smart-extraction-for-scraping/1562/2 "2025-03-05T20:26:18Z")

</div>

This topic was automatically closed 4 days after the last reply. New replies are no longer allowed.
