# Crawling outside the domain

**URL:** <https://forum.gumloop.com/t/crawling-outside-the-domain/1530>\
**Category:** Get Help\
**Tags:** Website-Crawler\
**Created:** [February 28, 2025, 5:39am UTC](https://forum.gumloop.com/t/crawling-outside-the-domain/1530 "2025-02-28T05:39:14Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![shinobi](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.gumloop.com/shinobi/32/455_2.png) [@shinobi](https://forum.gumloop.com/u/shinobi)\
**Post date:** [February 28, 2025, 5:39am UTC](https://forum.gumloop.com/t/crawling-outside-the-domain/1530/1 "2025-02-28T05:39:14Z")

</div>

![Screenshot 2025-02-28 at 9.37.48 AM](https://canada1.discourse-cdn.com/flex027/uploads/gumloop/original/1X/cea4d9335ce6b71141554c912955219f900d88f5.png)  
Even though I made sure the node is only run for same domain,  
the Run log showed it is crawling other URLs, and it was taking too long even for a single depth.

---

<div class="post-metadata">

**Author:** ![Wasay-Gumloop](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.gumloop.com/wasay-gumloop/32/69_2.png) [@Wasay-Gumloop](https://forum.gumloop.com/u/Wasay-Gumloop)\
**Post date:** [February 28, 2025, 10:48pm UTC](https://forum.gumloop.com/t/crawling-outside-the-domain/1530/3 "2025-02-28T22:48:08Z")

</div>

Hey @shinobi - Can you share the run link from the [https://www.gumloop.com/history](https://www.gumloop.com/history) page please? Please also set the share access to ‘anyone with the link can view’ under the share button.

 ![image](https://canada1.discourse-cdn.com/flex027/uploads/gumloop/original/1X/10f347a006c93b15f601badb27eaf7f6e69a8f39.png)

---

<div class="post-metadata">

**Author:** ![Wasay-Gumloop](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.gumloop.com/wasay-gumloop/32/69_2.png) [@Wasay-Gumloop](https://forum.gumloop.com/u/Wasay-Gumloop)\
**Post date:** [February 28, 2025, 10:50pm UTC](https://forum.gumloop.com/t/crawling-outside-the-domain/1530/4 "2025-02-28T22:50:24Z")

</div>

As for speed, the `Website Crawler` is more in depth hence slower but the `Web Agent Scraper` with the action `Get all URLs` is faster. It outputs the file URLs separated by a comma so you can use a `Split Text` node to get a list of all the URLs.

Here’s an eg: [https://www.gumloop.com/pipeline?workbook\_id=qwHFcSusCrk7QMZNgotwND&run\_id=QwoDMWNhefEKzhJmQGEkbj](https://www.gumloop.com/pipeline?workbook_id=qwHFcSusCrk7QMZNgotwND&run_id=QwoDMWNhefEKzhJmQGEkbj)

---

<div class="post-metadata">

**Author:** ![shinobi](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.gumloop.com/shinobi/32/455_2.png) [@shinobi](https://forum.gumloop.com/u/shinobi)\
**Post date:** [March 1, 2025, 4:14am UTC](https://forum.gumloop.com/t/crawling-outside-the-domain/1530/5 "2025-03-01T04:14:53Z")

</div>

Thanks Wasay, actually I got confused by the “Use only Same domain” flag.  
It was actually turned off. I turned it on and it was working as expected.

---

<div class="post-metadata">

**Author:** ![shinobi](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.gumloop.com/shinobi/32/455_2.png) [@shinobi](https://forum.gumloop.com/u/shinobi)\
**Post date:** [March 1, 2025, 4:18am UTC](https://forum.gumloop.com/t/crawling-outside-the-domain/1530/6 "2025-03-01T04:18:03Z")

</div>

[https://www.gumloop.com/pipeline?workbook\_id=2F8UxyKQyMnLAbkmHNgRjp](https://www.gumloop.com/pipeline?workbook_id=2F8UxyKQyMnLAbkmHNgRjp)

since you asked, here is my workbook link anyway.

---

<div class="post-metadata">

**Author:** ![Wasay-Gumloop](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.gumloop.com/wasay-gumloop/32/69_2.png) [@Wasay-Gumloop](https://forum.gumloop.com/u/Wasay-Gumloop)\
**Post date:** [March 1, 2025, 5:37am UTC](https://forum.gumloop.com/t/crawling-outside-the-domain/1530/7 "2025-03-01T05:37:11Z")

</div>

> [@shinobi](#):
>
> Thanks Wasay, actually I got confused by the “Use only Same domain” flag.  
> It was actually turned off. I turned it on and it was working as expected.

Awesome, glad you were able to solve it!

---

<div class="post-metadata">

**Author:** ![Wasay-Gumloop](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.gumloop.com/wasay-gumloop/32/69_2.png) [@Wasay-Gumloop](https://forum.gumloop.com/u/Wasay-Gumloop)\
**Post date:** [March 1, 2025, 5:38am UTC](https://forum.gumloop.com/t/crawling-outside-the-domain/1530/8 "2025-03-01T05:38:34Z")

</div>

> [@shinobi](#):
>
> [https://www.gumloop.com/pipeline?workbook\_id=2F8UxyKQyMnLAbkmHNgRjp](https://www.gumloop.com/pipeline?workbook_id=2F8UxyKQyMnLAbkmHNgRjp)
> 
> since you asked, here is my workbook link anyway.

Would recommend looking into subflows + error shield to make this flow more robust: [https://docs.gumloop.com/core-concepts/subflows](https://docs.gumloop.com/core-concepts/subflows)

---

<div class="post-metadata">

**Author:** ![system](https://canada1.discourse-cdn.com/flex027/uploads/gumloop/original/1X/d02ff8efcde6ac5d3ac8aa0376137d5f804da549.png) [@system](https://forum.gumloop.com/u/system)\
**Post date:** [March 5, 2025, 5:39pm UTC](https://forum.gumloop.com/t/crawling-outside-the-domain/1530/9 "2025-03-05T17:39:14Z")

</div>

This topic was automatically closed 4 days after the last reply. New replies are no longer allowed.
