# Scraping pagination links

**URL:** <https://forum.gumloop.com/t/scraping-pagination-links/3157>\
**Category:** Get Help\
**Tags:** Website-Scraper\
**Created:** [July 9, 2025, 7:22pm UTC](https://forum.gumloop.com/t/scraping-pagination-links/3157 "2025-07-09T19:22:03Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![OrionSeven](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.gumloop.com/orionseven/32/1227_2.png) [@OrionSeven](https://forum.gumloop.com/u/OrionSeven)\
**Post date:** [July 9, 2025, 7:22pm UTC](https://forum.gumloop.com/t/scraping-pagination-links/3157/1 "2025-07-09T19:22:03Z")

</div>

I’m scraping a website where the pagination links are time based and you won’t know the link till you scrape the site. (a while loop)

So I’m starting with the first link, extracting the next page link on the home page, and then I have a subflow that takes loads the next page link, scrapes it, gets the next link, checks if the date in the link is newer than 7 days (I only want to scrape back 7 days), and returns a boolean if it’s newer and the next url.

But how can I loop on this? Gummie help told me to use an if-else and loop the subflow back on itself.

But then when it hits the if-else it fails. I get " Automation Failed!  
Could not find next node to run. Please double check the node connections."

I’ve connected the if back to the subflow, and added the if pagination\_url to the subflows url input. (it has two). But nothings seems to be working.

Here’s my shared workbook. [https://www.gumloop.com/pipeline?workbook\_id=xir9wbPRz5rqMgJwe8sn3S&run\_id=j3FzhvqD7ze4xNjWrEgLV5](https://www.gumloop.com/pipeline?workbook_id=xir9wbPRz5rqMgJwe8sn3S&run_id=j3FzhvqD7ze4xNjWrEgLV5)

Thanks!

---

<div class="post-metadata">

**Author:** ![Gumloop\_Bot](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.gumloop.com/gumloop_bot/32/504_2.png) [@Gumloop\_Bot](https://forum.gumloop.com/u/Gumloop_Bot)\
**Post date:** [July 9, 2025, 7:22pm UTC](https://forum.gumloop.com/t/scraping-pagination-links/3157/2 "2025-07-09T19:22:04Z")

</div>

Hey @OrionSeven! If you’re reporting an issue with a flow or an error in a run, please include the **run link** and make sure it’s **shareable** so we can take a look.

1. **Find your run link** on the history page. Format: `https://www.gumloop.com/pipeline?run_id={{your_run_id}}&workbook_id={{workbook_id}}`

2. **Make it shareable** by clicking **“Share”** → ‘Anyone with the link can view’ in the top-left corner of the flow screen.  
 ![GIF guide](https://canada1.discourse-cdn.com/flex027/uploads/gumloop/original/1X/92c00c612fd65b861e706dd1b5eec5ec61677835.gif)

3. **Provide details** about the issue—more context helps us troubleshoot faster.

You can find your run history here: [https://www.gumloop.com/history](https://www.gumloop.com/history)

---

<div class="post-metadata">

**Author:** ![Wasay-Gumloop](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.gumloop.com/wasay-gumloop/32/69_2.png) [@Wasay-Gumloop](https://forum.gumloop.com/u/Wasay-Gumloop)\
**Post date:** [July 10, 2025, 4:32am UTC](https://forum.gumloop.com/t/scraping-pagination-links/3157/3 "2025-07-10T04:32:45Z")

</div>

Hey @OrionSeven – Just to clarify, what you’re trying to set up would be a circular reference, where it loops through the same subflow – is that right? From what I can see, you’re looking to create a setup that references itself. At the moment, circular references aren’t possible in Gumloop. The main reason is that a circular reference could potentially run infinitely, which makes it difficult to control or stop, and it can get pretty confusing quickly.

Since the [Hacker News website](https://thehackernews.com/search?updated-max=2025-06-26T14:15:00%2B05:30&max-results=12&start=48&by-date=false) has a specific pagination URL format, each time you click the next page button, it basically adds results to the start page, and there are some initial links as well. What you can do is copy this format and have Gummie or AI write up a script that can output a list of paginated URLs for Hacker News, you can use this through the Run Code node. For example, if you want 2 paginated URLs or 3 paginated URLs, it outputs a list that you can simply loop over with your website scraper node. Would something like that work for you?

---

<div class="post-metadata">

**Author:** ![OrionSeven](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.gumloop.com/orionseven/32/1227_2.png) [@OrionSeven](https://forum.gumloop.com/u/OrionSeven)\
**Post date:** [July 10, 2025, 4:25pm UTC](https://forum.gumloop.com/t/scraping-pagination-links/3157/4 "2025-07-10T16:25:50Z")

</div>

It’s [thehackernews.com](http://thehackernews.com) not the YC site. And their links are not built out in the same manner thus the need for the loop.

I’ll try something else. Thanks for the response.

---

<div class="post-metadata">

**Author:** ![Wasay-Gumloop](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.gumloop.com/wasay-gumloop/32/69_2.png) [@Wasay-Gumloop](https://forum.gumloop.com/u/Wasay-Gumloop)\
**Post date:** [July 12, 2025, 2:18am UTC](https://forum.gumloop.com/t/scraping-pagination-links/3157/5 "2025-07-12T02:18:02Z")

</div>

No problem! I was actually referring to [thehackernews.com](http://thehackernews.com) as well. If you click on `Next Page` it starts to add query parameters in the URL, example: `https://thehackernews.com/search?updated-max=2025-06-26T14:15:00%2B05:30&max-results=12&start=48&by-date=false`

The idea was to use this format in a `Run Code` node and have it output a `List` of paginated URLs that you can loop through the `Website Scraper` node. Essentially, only `max-results=12&start=48` are changing in the URL along with the date.

---

<div class="post-metadata">

**Author:** ![system](https://canada1.discourse-cdn.com/flex027/uploads/gumloop/original/1X/d02ff8efcde6ac5d3ac8aa0376137d5f804da549.png) [@system](https://forum.gumloop.com/u/system)\
**Post date:** [August 1, 2025, 2:18am UTC](https://forum.gumloop.com/t/scraping-pagination-links/3157/6 "2025-08-01T02:18:19Z")

</div>

This topic was automatically closed 20 days after the last reply. New replies are no longer allowed.
