# Website Scraper Failures when Looping, but perfect when individual flows are run

**URL:** <https://forum.gumloop.com/t/website-scraper-failures-when-looping-but-perfect-when-individual-flows-are-run/2739>\
**Category:** Get Help\
**Tags:** Website-Scraper\
**Created:** [May 14, 2025, 3:59am UTC](https://forum.gumloop.com/t/website-scraper-failures-when-looping-but-perfect-when-individual-flows-are-run/2739 "2025-05-14T03:59:47Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![BigGummiePlans](https://avatars.discourse-cdn.com/v4/letter/b/4491bb/32.png) [@BigGummiePlans](https://forum.gumloop.com/u/BigGummiePlans)\
**Post date:** [May 14, 2025, 3:59am UTC](https://forum.gumloop.com/t/website-scraper-failures-when-looping-but-perfect-when-individual-flows-are-run/2739/1 "2025-05-14T03:59:47Z")

</div>

I have a flow to take a URL, scrape data, AI analyze, and output back into sheets.  
When the URL are input one-at-a-time, manually into the flow, everything works perfectly.  
When the URL’s are read from google sheet, some scrapes fail. Repeated runs give different URL failures, so it feels random.  
I increased timeout on website scraper and wouldn’t say that changed the result all that much.  
I reverted back to individual input, and again, it works fine.

Pipeline of Individual Flow - Perfect Run.  
[https://www.gumloop.com/pipeline?run\_id=gNrnJTKez78s9GzHg9aJ6U&workbook\_id=dSV9q7fDmaz4jZTED6WnJj&tab=1](https://www.gumloop.com/pipeline?run_id=gNrnJTKez78s9GzHg9aJ6U&workbook_id=dSV9q7fDmaz4jZTED6WnJj&tab=1)

Pipeline of Looped Flow - some URL scrapes fail.  
[https://www.gumloop.com/pipeline?run\_id=RonNduSJMe2A5pBPapHn3a&workbook\_id=dSV9q7fDmaz4jZTED6WnJj](https://www.gumloop.com/pipeline?run_id=RonNduSJMe2A5pBPapHn3a&workbook_id=dSV9q7fDmaz4jZTED6WnJj)

Here is sheets with input URLS and expected results (shared permissions lifted)

> **[PLAN URL](https://docs.google.com/spreadsheets/d/1uXTddu624X4fwToT1QkDgq6f3XlHz1I-QBi6nCb5zbU/edit?usp=sharing)**
>
> This Sheet is private

---

<div class="post-metadata">

**Author:** ![Gumloop\_Bot](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.gumloop.com/gumloop_bot/32/504_2.png) [@Gumloop\_Bot](https://forum.gumloop.com/u/Gumloop_Bot)\
**Post date:** [May 14, 2025, 3:59am UTC](https://forum.gumloop.com/t/website-scraper-failures-when-looping-but-perfect-when-individual-flows-are-run/2739/2 "2025-05-14T03:59:48Z")

</div>

Hey @BigGummiePlans! If you’re reporting an issue with a flow or an error in a run, please include the **run link** and make sure it’s **shareable** so we can take a look.

1. **Find your run link** on the history page. Format: `https://www.gumloop.com/pipeline?run_id={{your_run_id}}&workbook_id={{workbook_id}}`

2. **Make it shareable** by clicking **“Share”** → ‘Anyone with the link can view’ in the top-left corner of the flow screen.  
 ![GIF guide](https://canada1.discourse-cdn.com/flex027/uploads/gumloop/original/1X/92c00c612fd65b861e706dd1b5eec5ec61677835.gif)

3. **Provide details** about the issue—more context helps us troubleshoot faster.

You can find your run history here: [https://www.gumloop.com/history](https://www.gumloop.com/history)

---

<div class="post-metadata">

**Author:** ![Wasay-Gumloop](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.gumloop.com/wasay-gumloop/32/69_2.png) [@Wasay-Gumloop](https://forum.gumloop.com/u/Wasay-Gumloop)\
**Post date:** [May 15, 2025, 1:37am UTC](https://forum.gumloop.com/t/website-scraper-failures-when-looping-but-perfect-when-individual-flows-are-run/2739/3 "2025-05-15T01:37:53Z")

</div>

Hey @BigGummiePlans –

If you click on the subflow runs that failed on the main flow where URLs are being pulled from the `Sheet Reader` node you’ll notice that on each failed run the `Duplicate` node failed, this is because the `Extract Data` node could not extract relevant data for a specific field and output a blank string, in which case the `Duplicate` node did not have any input to duplicate.

View failed subflow run:

 ![image](https://canada1.discourse-cdn.com/flex027/uploads/gumloop/original/2X/c/c5535da1376fc78fcbddfe99b36c6f09e34c3b3b.png)

Subflow error:

 ![Screenshot 2025-05-15 at 6.28.37 AM](https://canada1.discourse-cdn.com/flex027/uploads/gumloop/original/2X/7/7e5744b0e581b8e7c4672cc754a65a566457894d.png)

I’m not exactly sure why a `Duplicate` node is needed here, I’m assuming it was added to meet the `List` input when `Writer Mode` on the sheet writer node is set to `Write to Column`. A quick and easy solution would be to simply remove the duplicate nodes and change the writer mode to `Add a Single New Row` – this will allow you to write empty/blank strings:

 ![Screenshot 2025-05-15 at 6.32.37 AM](https://canada1.discourse-cdn.com/flex027/uploads/gumloop/original/2X/f/fc267e6d26a54d24a156b1e8eea937c19e486579.png)

A more robust solution here though would involve two edits:

1. In your subflow you should use a `Google Sheet Updater` node to find and update the existing row since you’re using a single spreadsheet to read and write. [More on Sheet Updater node here.](https://docs.gumloop.com/nodes/integrations/gsheets_updater)

2. In your main flow you should wrap the subflow in an `Error Shield`, that way if anything fails for one off URLs they can be skipped without halting the entire flow

Eg setup: [https://www.gumloop.com/pipeline?workbook\_id=8ruwLErkGcyUf7XbVmWiEZ](https://www.gumloop.com/pipeline?workbook_id=8ruwLErkGcyUf7XbVmWiEZ)

Let me know if this makes sense and works for you 🙂

---

<div class="post-metadata">

**Author:** ![BigGummiePlans](https://avatars.discourse-cdn.com/v4/letter/b/4491bb/32.png) [@BigGummiePlans](https://forum.gumloop.com/u/BigGummiePlans)\
**Post date:** [May 15, 2025, 6:46am UTC](https://forum.gumloop.com/t/website-scraper-failures-when-looping-but-perfect-when-individual-flows-are-run/2739/4 "2025-05-15T06:46:59Z")

</div>

Thanks Wasay.

First recommendation was correct, removed duplicate, all runs execute.  
However, the real underlying issue is unsolved.  
Recapping :  
I have URL as inputs, to be scraped.  
If I manually input each URL one-at-a-time into my sub flow - each runs successfully, and returns results.  
When i enter the same URLS as a list, and use primary flow - the runs complete, but website scrapes randomly collect nothing.

Here is the output list, showing the results of the extracts in both individual runs, list runs (x2).

From looking at the runs, website scraper seems to have collected nothing in some instances.

I have adjusted the time out, but it seems to not help.

 ![Screenshot 2025-05-15 at 16.34.29](https://canada1.discourse-cdn.com/flex027/uploads/gumloop/original/2X/6/69efe937081585ccb132325063d6611f3a77bc5a.png)

Workbook (Shared) : [https://www.gumloop.com/pipeline?workbook\_id=dSV9q7fDmaz4jZTED6WnJj&run\_id=W9nL6M3LPc7jAEBxezfWKC&tab=1](https://www.gumloop.com/pipeline?workbook_id=dSV9q7fDmaz4jZTED6WnJj&run_id=W9nL6M3LPc7jAEBxezfWKC&tab=1)

Results

 ![Screenshot 2025-05-15 at 16.47.54](https://canada1.discourse-cdn.com/flex027/uploads/gumloop/original/2X/7/75fc21662f42560ba7f246d5ae84df59a6cfe595.png)

One more thing, I just did a input list, as a trigger, new row add, and added each URL about 2 mins apart and got a perfect result. See sheets in screen shot, link here [PLAN URL - Google Sheets](https://docs.google.com/spreadsheets/d/1uXTddu624X4fwToT1QkDgq6f3XlHz1I-QBi6nCb5zbU/edit?gid=0#gid=0)

me thinks there is an issue with parallel processing website scrapes at the same time, perhaps its memory overhead, container time.

---

<div class="post-metadata">

**Author:** ![Wasay-Gumloop](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.gumloop.com/wasay-gumloop/32/69_2.png) [@Wasay-Gumloop](https://forum.gumloop.com/u/Wasay-Gumloop)\
**Post date:** [May 16, 2025, 1:08am UTC](https://forum.gumloop.com/t/website-scraper-failures-when-looping-but-perfect-when-individual-flows-are-run/2739/5 "2025-05-16T01:08:49Z")

</div>

Appreciate the detailed context. I think you’re right, the site seems to be blocking the scraper. I’ll flag this with the scraper provider.

---

<div class="post-metadata">

**Author:** ![system](https://canada1.discourse-cdn.com/flex027/uploads/gumloop/original/1X/d02ff8efcde6ac5d3ac8aa0376137d5f804da549.png) [@system](https://forum.gumloop.com/u/system)\
**Post date:** [May 21, 2025, 1:08am UTC](https://forum.gumloop.com/t/website-scraper-failures-when-looping-but-perfect-when-individual-flows-are-run/2739/6 "2025-05-21T01:08:58Z")

</div>

This topic was automatically closed 5 days after the last reply. New replies are no longer allowed.
