Proxies That Work logo

How to Crawl a Website Without Getting Blocked: 15 Practical Tips for 2026

By Jesse Lewis•8/31/2026•5 min read

Website crawling becomes unreliable when request volume, concurrency, retries, and session behavior are poorly controlled.

A block is rarely caused by one isolated request. More often, problems appear when several operational signals accumulate at the same time: excessive concurrency, repeated requests, unstable sessions, aggressive retries, or traffic that exceeds what the target can reasonably handle.

The goal of a production crawler should not be to eliminate every possible block. It should be to collect permitted data predictably, respond correctly to server feedback, and reduce unnecessary load on both your infrastructure and the target website.

These 15 practices can help improve crawl reliability.

1. Define a Crawl Budget Before You Start

Before sending requests, determine exactly what the crawler needs to accomplish.

Define:

  • the pages or data you actually need;
  • acceptable requests per minute for each domain;
  • maximum concurrency;
  • acceptable error rate;
  • required crawl completion time;
  • refresh frequency.

A crawler without explicit limits can generate substantially more traffic than the task requires.

For example, if only 50,000 product pages need daily updates, there may be no reason to recrawl category pages, static assets, and unchanged URLs on every run.

A crawl budget turns an open-ended scraping process into a controlled workload.

2. Control Concurrency, Not Just Request Delays

Adding delays between requests can help, but concurrency is equally important.

A crawler might average only a few requests per second while still opening dozens of connections simultaneously.

Instead of relying only on sleep intervals, set limits for:

  • concurrent requests per domain;
  • concurrent browser sessions;
  • concurrent requests per proxy;
  • concurrent workers.

A simple rule such as:

Domain → Maximum N active requests

is often more effective than adding random delays everywhere.

Concurrency limits also make crawl performance easier to predict.

3. Treat HTTP Errors as Operational Feedback

HTTP responses provide information about how the target is handling your traffic.

Important examples include:

  • 403 Forbidden: access is being denied;
  • 429 Too Many Requests: the server is rate limiting requests;
  • 503 Service Unavailable: the service may be overloaded or temporarily unavailable.

Do not respond to every error by immediately retrying.

Instead:

  1. classify the response;
  2. check whether failures are isolated or increasing;
  3. reduce concurrency if necessary;
  4. apply appropriate backoff;
  5. determine whether the problem is target-wide, session-specific, or proxy-specific.

A structured scraper-block debugging process can help distinguish network failures, rate limits, application errors, and target-side access restrictions before changes are made.

4. Rotate Proxies According to the Workload

Changing IP addresses more frequently is not automatically better.

Proxy rotation should support the workload rather than introduce unnecessary session changes.

For stateless requests, rotation across a healthy proxy pool may distribute traffic effectively.

For stateful interactions, maintaining the same proxy for a defined session can be more appropriate.

Useful strategies include:

  • round-robin allocation;
  • health-aware selection;
  • sticky sessions;
  • per-domain pools;
  • cooldown periods for degraded proxies.

If your application controls the routing directly, a structured Python proxy rotation implementation can keep selection, health monitoring, retries, and application logic separate.

5. Separate Discovery Crawls From Extraction Crawls

Discovery and extraction solve different problems.

Discovery Crawling

Discovery identifies:

  • new URLs;
  • changed sections;
  • pagination;
  • sitemaps;
  • category structures.

It can usually operate with lightweight requests and relatively infrequent refreshes.

Extraction Crawling

Extraction retrieves the specific pages required for structured data.

These requests may need stricter:

  • scheduling;
  • validation;
  • retry logic;
  • session management.

Separating the two workloads prevents discovery activity from consuming capacity needed for high-value extraction jobs.

It also makes each stage easier to monitor.

6. Cache and Deduplicate Requests

Repeatedly requesting information you already have wastes bandwidth, proxy capacity, and target-server resources.

Implement mechanisms such as:

  • normalized URL storage;
  • seen-URL tracking;
  • content caching;
  • redirect caching;
  • duplicate-job detection;
  • conditional requests using ETag or Last-Modified where supported.

For example, if a crawler encounters:

example.com/product?id=100
example.com/product?id=100&utm_source=test

URL normalization may determine that both references point to the same underlying page.

Avoiding unnecessary duplicate requests can materially reduce crawl volume.

7. Keep Request Configuration Consistent

Production crawlers should use stable and deliberate HTTP configuration.

This can include:

  • a documented User-Agent;
  • appropriate Accept headers;
  • consistent language preferences where required;
  • stable session configuration;
  • predictable cookie handling.

Avoid changing request configuration randomly between otherwise related requests.

Consistency also improves debugging because developers can reproduce how the crawler communicated with the target.

Where a website publishes crawler guidance or requires a specific User-Agent format, follow those requirements.

8. Use the Simplest Tool That Works

Not every website requires browser automation.

For primarily server-rendered pages, an HTTP client is usually:

  • faster;
  • less resource-intensive;
  • easier to scale;
  • easier to debug.

Browser automation becomes more appropriate when required content depends on:

  • client-side JavaScript;
  • dynamic application state;
  • browser-specific interaction;
  • authenticated multi-step flows.

The comparison between headless browsers and HTTP clients can help determine which approach matches a particular target.

Using a browser for every URL when simple HTTP requests would work can dramatically increase infrastructure cost and crawl complexity.

9. Preserve Sessions When the Workflow Requires Them

Some workflows depend on continuity between requests.

Examples include:

  • authenticated sessions;
  • shopping carts;
  • pagination using session state;
  • applications that issue session cookies;
  • multi-step forms.

In those situations, changing IPs or creating new sessions between every request can break application state.

A better approach may be:

Create session → Assign proxy → Complete related requests → Close session

Session persistence should be used because the target workflow requires it, not as a universal crawler setting.

10. Use Predictable, Conservative Scheduling

Request timing should be controlled rather than generated by an unrestricted loop.

Useful controls include:

  • maximum request rate;
  • concurrency ceilings;
  • scheduled crawl windows;
  • per-domain limits;
  • queue-based dispatching.

Small amounts of timing variation may naturally occur in distributed systems, but deliberate request-rate controls are more important than artificial randomness.

The objective is simple: avoid unnecessary traffic spikes.

11. Back Off When You Receive Rate-Limit Signals

HTTP 429 responses are explicit signals that request volume should be reduced.

When the server supplies a Retry-After header, respect it where applicable.

A retry strategy can include:

  1. pause the affected request;
  2. reduce concurrency;
  3. increase the interval before retrying;
  4. reschedule the request;
  5. monitor whether the error rate returns to normal.

Exponential backoff is commonly useful for temporary failures.

For example:

Retry 1 → wait 2 seconds
Retry 2 → wait 4 seconds
Retry 3 → wait 8 seconds
Retry 4 → wait 16 seconds

Place an upper limit on retries so that persistent failures do not create infinite loops.

CAPTCHA or explicit access-challenge pages should also be treated as signals to reassess whether automated access is appropriate rather than repeatedly forcing the same request.

12. Monitor Proxy and IP Performance Over Time

Proxy performance changes.

An IP that works reliably today may later experience:

  • higher latency;
  • more connection failures;
  • increased 403 responses;
  • increased 429 responses;
  • target-specific access problems.

Track metrics by individual IP, subnet, pool, and target domain.

Useful measurements include:

  • request success percentage;
  • 403 rate;
  • 429 rate;
  • timeout frequency;
  • latency percentiles;
  • consecutive failures;
  • last successful request.

A broader IP reputation management strategy can help teams decide when to place degraded addresses in cooldown or remove them from a production pool.

Technical capability does not automatically mean a crawl should proceed.

Before operating a production crawler, review:

  • robots.txt;
  • website terms and policies;
  • authentication requirements;
  • applicable privacy obligations;
  • contractual restrictions;
  • relevant laws and regulations.

Organizations should also distinguish between publicly accessible information and information protected by authentication or other access controls.

For teams using proxies at scale, documented bulk proxy compliance practices can help integrate technical controls with legal, privacy, and governance requirements.

14. Design for Graceful Degradation

A reliable crawler should continue delivering useful results when conditions deteriorate.

Instead of repeatedly forcing the original crawl plan, the system can adapt.

Examples include:

  • reducing concurrency;
  • lowering crawl frequency;
  • prioritizing high-value URLs;
  • postponing lower-priority jobs;
  • switching from full collection to sampling;
  • pausing affected targets;
  • resuming from checkpoints.

Suppose a crawler normally processes one million URLs per day.

If success rates suddenly decline, it may be better to collect the 100,000 most important URLs reliably than continue retrying the entire million-URL workload.

Graceful degradation protects both infrastructure capacity and data quality.

15. Test Small Before Scaling

Blocking and performance problems become substantially more expensive once a crawler reaches production volume.

Before scaling:

  1. run a limited sample;
  2. measure success rate;
  3. record latency;
  4. inspect HTTP response distribution;
  5. verify extracted data;
  6. gradually increase concurrency;
  7. monitor how performance changes.

For example:

1 worker
   ↓
5 workers
   ↓
10 workers
   ↓
25 workers
   ↓
Production capacity

At each stage, compare:

  • request success rate;
  • 403 and 429 frequency;
  • latency;
  • retry volume;
  • data completeness.

Treat crawler deployment like infrastructure scaling rather than launching an unrestricted script.

Metrics Worth Monitoring

A production crawler should expose enough telemetry to determine whether performance is improving or degrading.

Useful metrics include:

Metric Why It Matters
Request success rate Measures overall crawl reliability
HTTP 403 rate Highlights access-denied responses
HTTP 429 rate Identifies rate limiting
HTTP 5xx rate Identifies target-side/server errors
Timeout rate Shows network or capacity problems
p95 latency Detects slower requests hidden by averages
Retry rate Reveals inefficient or unstable jobs
Crawl completion rate Measures whether scheduled work finishes
Data completeness Confirms expected records were actually collected
Requests per proxy Helps identify traffic concentration

Network metrics should always be connected with data-quality metrics.

A crawler returning HTTP 200 responses is not useful if the expected information is missing or invalid.

A Practical Crawl-Control Workflow

A resilient crawling pipeline can follow this pattern:

URL Queue
   ↓
Domain Scheduler
   ↓
Concurrency / Rate Limits
   ↓
Proxy Allocator
   ↓
HTTP Client or Browser
   ↓
Response Classification
   ↓
Data Validation
   ↓
Storage
   ↓
Metrics + Feedback

The feedback layer can then adjust:

  • crawl frequency;
  • concurrency;
  • retries;
  • proxy allocation;
  • job priority.

This creates an adaptive system without relying on uncontrolled request behavior.

Frequently Asked Questions

Why do websites block crawlers?

Websites may restrict automated traffic for many reasons, including server protection, rate limiting, security policies, terms of service, or abuse prevention. High concurrency, repeated requests, and excessive retries can increase the likelihood of restrictions.

Does using more proxies prevent blocking?

No. A larger proxy pool provides more traffic-distribution capacity, but poor request behavior can still generate failures. Pool size should be combined with rate control, health monitoring, retries, and target-specific scheduling.

What should a crawler do after receiving HTTP 429?

Reduce the request rate and respect Retry-After instructions when provided. Repeatedly retrying immediately can worsen rate limiting.

Is HTTP 403 always caused by the proxy?

No. A 403 can result from access policies, authentication problems, application rules, geographic restrictions, or other factors. Diagnose the response before assuming that changing proxies will solve it.

Should I use a headless browser for scraping?

Only when the target requires browser capabilities such as JavaScript rendering or interactive workflows. For server-rendered pages, a normal HTTP client is generally simpler and more efficient.

How much concurrency should a crawler use?

There is no universal number. Start conservatively, measure target response and crawler performance, and increase concurrency gradually while monitoring errors, latency, and data completeness.

How can I reduce unnecessary crawl traffic?

Use URL deduplication, caching, incremental updates, conditional requests, crawl prioritization, and separate discovery from extraction.

Final Thoughts

Reliable crawling is less about finding a single configuration that avoids every block and more about building infrastructure that controls traffic, responds to server feedback, and adapts when conditions change.

The most effective crawlers combine:

  • explicit crawl budgets;
  • concurrency controls;
  • responsible request pacing;
  • workload-aware proxy allocation;
  • session management;
  • caching and deduplication;
  • error-specific backoff;
  • data validation;
  • monitoring;
  • graceful degradation.

When these controls are in place, large crawls become easier to operate, troubleshoot, and scale without continuously rebuilding the collection system.

About the Author

J

Jesse Lewis

Jesse Lewis is a researcher and content contributor for ProxiesThatWork, covering compliance trends, data governance, and the evolving relationship between AI and proxy technologies. He focuses on helping businesses stay compliant while deploying efficient, scalable data-collection pipelines.

Proxies That Work logo
© 2026 ProxiesThatWork LLC. All Rights Reserved.