For HighLevel's public API 2.0 endpoints authenticated with OAuth, the current published rate limits are:
- 100 API requests per 10 seconds
- 200,000 API requests per day
Those limits are counted per Marketplace app, per resource, where the resource is a specific HighLevel Location or Company. An app installed on Location A and Location B therefore receives an independent API budget for each resource.
But those numbers should not be treated as a universal limit for every HighLevel API surface.
For example, HighLevel currently documents reduced limits for Sandbox Private Integration Tokens and a separate rate limit for the Agent Studio Public API.
That makes the real production question bigger than:
How many HighLevel API requests can I make?
A reliable integration also needs to know:
- which quota applies;
- which Location or Company owns that quota;
- how much of the current budget remains;
- whether a
429represents temporary burst pressure or exhausted daily capacity; - how concurrent workers share the same quota;
- when retrying is useful;
- and when retrying can create a different business problem.
The most important rule is simple:
Use the rate-limit state HighLevel returns instead of designing the integration around one hard-coded number.
What Are the Current GoHighLevel API Rate Limits?
For the documented public API 2.0 OAuth context, HighLevel currently publishes the following limits:
Limit
Allowance
Burst
100 requests per 10 seconds
Daily
200,000 requests per day
Scope
Per Marketplace app per Location or Company
HighLevel explicitly scopes those limits by app + resource.
For example:
Marketplace App
│
├── Location A
│ ↓
│ its own API budget
│
└── Location B
↓
separate API budgetIf the application uses its full burst allowance against Location A, that does not mean Location B has consumed the same quota.
But several processes calling the same app + Location A are still operating against the same resource budget.
Does 100 requests per 10 seconds mean exactly 10 requests per second?
No.
Mathematically, 100 requests over 10 seconds has an average of 10 requests per second, but HighLevel documents this as a burst limit over an interval.
The documentation does not tell us to model the system as exactly 10 requests in every individual second, nor does it require developers to reverse-engineer the internal rate-limiting algorithm.
Instead, HighLevel exposes the current quota state in API response headers and says those headers should be treated as authoritative.
So avoid turning:
100 requests / 10 secondsinto an undocumented rule such as:
Exactly 10 requests every secondYour integration should react to the API state HighLevel actually returns.
Which Rate Limit Applies to My HighLevel Integration?
Before building any retry or throttling logic, identify the API surface and authentication context you are actually using.
The common mistake is to see 100 requests / 10 seconds in one HighLevel document and apply it everywhere.
Current official documentation already shows why that is unsafe.
API context
Current documented behavior
Public API 2.0 with OAuth
100 requests / 10 sec + 200,000/day
Sandbox PIT
25 requests / 10 sec + 10,000/day
Production PIT
Standard production limits; verify the applicable production context
Agent Studio Public API
300 requests/minute per sub-account
Other specialized API surfaces
Check that API's current documentation
HighLevel's Sandbox PIT documentation says Sandbox Private Integration Tokens are limited to 25 requests per 10 seconds and 10,000 requests per day. It also says creating multiple PITs does not multiply the sandbox allowance.
HighLevel separately documents 300 requests per minute per sub-account across its Agent Studio Public API endpoints. (HighLevel Support Portal)
Therefore:
“The GoHighLevel API limit is always 100 requests every 10 seconds” is too broad.
A better process is:
Which API am I calling?
↓
Which authentication/environment?
↓
What does its current documentation say?
↓
What quota headers did the actual response return?That is especially important for integrations that use several HighLevel products or move from Sandbox into production.
Read the X-RateLimit Headers Instead of Counting Requests Yourself
For the documented public API context, HighLevel returns rate-limit information on API responses.
The current documented headers are:
Header
Meaning
X-RateLimit-Limit-Daily
Daily request allowance
X-RateLimit-Daily-Remaining
Requests remaining today
X-RateLimit-Interval-Milliseconds
Length of the burst interval
X-RateLimit-Max
Maximum requests allowed in that interval
X-RateLimit-Remaining
Requests remaining in the current interval
HighLevel specifically recommends reading these values instead of counting requests independently because the response headers are authoritative.
Conceptually:
HighLevel API response
↓
X-RateLimit-Max
X-RateLimit-Remaining
X-RateLimit-Daily-Remaining
X-RateLimit-Interval-Milliseconds
↓
Update pacing / scheduling decisionThis matters because local counters can become inaccurate when:
- several application instances are running;
- multiple workers are processing the same Location;
- a deployment restarts;
- another service uses the same app/resource context;
- the documented quota changes.
A production integration should not assume an old blog post knows more about its current quota than the API response itself.

A 429 Can Represent Two Very Different Problems
429 Too Many Requests tells you that the current request exceeded an applicable rate limit.
But the correct recovery depends on which limit was exhausted.
Case 1: The burst allowance is exhausted
Imagine a bulk process sends a large number of calls against the same Location in a short period:
500 jobs
↓
many workers
↓
same HighLevel Location
↓
burst budget exhausted
↓
429If daily capacity is still available, this is primarily a pacing problem.
The work may be able to continue after the burst interval allows additional requests.
Case 2: The daily allowance is exhausted
Now consider:
X-RateLimit-Daily-Remaining = 0Waiting one second, five seconds, or another burst interval does not restore daily capacity.
Continuing to immediately retry only creates more failed work.
The integration instead needs to:
stop immediate retries
↓
defer remaining jobs
↓
retain their state
↓
resume when capacity is availableHighLevel's current official JavaScript/TypeScript SDK demonstrates exactly this conceptual distinction. When its optional rate-limit retry feature is enabled, a burst 429 can wait using the reported interval and retry, while a response showing zero daily requests remaining is surfaced immediately because another immediate attempt cannot succeed.
A useful decision tree is:
429 Too Many Requests
↓
Inspect current rate-limit state
↓
Daily remaining = 0?
/ \
Yes No
↓ ↓
Stop immediate Burst pressure
retry ↓
↓ wait / pace
Defer work ↓
bounded retryTreating every 429 as:
“sleep for a few seconds and try again”
misses this distinction.
Do Not Assume Every HighLevel 429 Includes Retry-After
Many generic HTTP retry examples tell developers to read the Retry-After header.
That is useful when an API provides it, but HighLevel's main Rate Limits documentation does not define Retry-After as the universal mechanism for API 2.0 rate-limit handling.
Instead, it documents the X-RateLimit-* headers.
The current official HighLevel SDK follows the same pattern.
For a burst 429, it uses the reported X-RateLimit-Interval-Milliseconds value. If rate-limit headers are unavailable, the SDK falls back to exponential backoff with jitter.
Therefore:
Respect Retry-After when the specific API you are calling returns and documents it, but do not assume every HighLevel 429 provides that header.Use the quota information the actual response gives you.
Prevent 429 Errors Instead of Only Reacting to Them
A production integration should not rely exclusively on this pattern:
send as quickly as possible
↓
hit 429
↓
wait
↓
retryThat is reactive throttling.
A better high-volume architecture controls the rate before the API needs to reject traffic.
For example:
Incoming jobs
↓
Queue
↓
Rate-aware scheduling
↓
Bounded workers
↓
HighLevel APIThis can reduce:
- unnecessary
429s; - retry storms;
- worker time spent sleeping;
- unpredictable latency;
- noisy error logs.
A small integration does not automatically need Redis, distributed workers, or a large queueing platform.
For a workflow making only a few calls, a simple rate limiter may be enough.
The architecture should match the workload.
The principle is:
Pace the work before concurrency turns a normal job into an API burst.
Concurrency and Request Rate Are Not the Same Thing
Two related controls are often confused.
Concurrency
How many API operations are executing at the same time?
Rate
How many requests begin during a given time window?
Those are not identical.
For example, 20 workers do not always produce exactly 20 requests per second. API calls take different amounts of time.
But high concurrency can produce a large burst when requests complete quickly.
Consider:
500 contact updates
↓
start everything at once
↓
many requests compressed
into a short periodA more controlled model is:
500 jobs
↓
bounded concurrency
+
rate pacing
↓
HighLevel APIThis matters with:
- queue workers;
- n8n item processing;
- parallel batch jobs;
- background services;
- large JavaScript
Promise.all()operations.
The goal is not to eliminate concurrency.
It is to keep concurrency from violating the API budget that applies to the resource being processed.
Multiple Workers Calling the Same Location Share the Same Quota
HighLevel scopes its documented OAuth quota by Marketplace app + resource.
It does not describe a separate quota for every application server, container, worker, or process.
That creates an important production consideration.
Suppose three workers all process the same HighLevel Location:
Worker 1 ─┐
Worker 2 ─┼──→ App X + Location A
Worker 3 ─┘From HighLevel's perspective, that is still traffic against the same app/resource budget.
If each worker independently assumes it owns the entire allowance, the combined traffic can exceed the actual quota.
For larger systems, rate control may therefore need to be coordinated across workers.
Conceptually:
Worker 1 ─┐
Worker 2 ─┼→ shared scheduling / quota coordination
Worker 3 ─┘
↓
App X + Location A
↓
HighLevel APIThat coordination might be implemented with a central queue, shared limiter, scheduler, or another mechanism appropriate to the application.
The implementation is less important than the principle:
Do not give every worker the illusion that it owns a quota HighLevel actually scopes to their shared resource.
Multi-Location Apps Need Per-Location Rate-Limit Thinking
The same scoping rule also creates an advantage for multi-location applications.
For the documented OAuth context:
Marketplace App
│
├── Location A → quota A
│
├── Location B → quota B
│
└── Location C → quota CHighLevel explicitly documents that installing the same app on multiple sub-accounts gives each installation an independent budget rather than dividing one global allowance among all of them.
That means a noisy Location should not automatically force every other Location to stop processing.
A useful scheduling model is:
Incoming jobs
↓
identify locationId
↓
resource-specific queue/pacing
↓
HighLevelIf Location A is temporarily throttled:
Location A
→ slow / deferwhile:
Location B
→ continue normallyprovided your own infrastructure does not unnecessarily couple them.
This is another reason tenant context needs to remain explicit throughout a multi-location integration.
Daily API Capacity Is a Budget Problem, Not a Backoff Problem
Burst throttling asks:
How quickly can this work be sent?
Daily capacity asks:
Should this process make this many calls at all?
Imagine a job that requires:
5,000 records
×
50 API requests per record
=
250,000 API callsIn the documented OAuth context, that exceeds a 200,000-request daily allowance for one app/resource even if the requests are paced perfectly. (App Marketplace)
No clever backoff algorithm changes the total number of requests.
The useful question becomes:
Why are we making 50 calls for each record?
Common opportunities to reduce usage include:
Pattern
Better question
Polling repeatedly
Can an event/webhook tell us when something changes?
Re-reading stable data
Can the result be cached?
Same lookup for every row
Can the value be reused?
Full sync every run
Can only changed records be processed?
One huge migration
Can the workload be scheduled across safe windows?
Avoid assuming a bulk endpoint exists unless the specific HighLevel API resource documents one.
The goal is not merely to send an inefficient integration more slowly.
It is to reduce unnecessary API work.
Use Webhooks When They Can Replace Wasteful Polling
Polling can consume API budget even when nothing has changed.
For example:
Every minute
↓
Fetch records
↓
Compare state
↓
Nothing changedRepeated across many Locations, this can become expensive in request count.
When HighLevel exposes an appropriate event, an event-driven design may be more efficient:
HighLevel event
↓
Webhook
↓
Process the actual changeThis does not mean every API read should become a webhook.
It means polling should be justified when an appropriate event-driven boundary exists.
Cache Data That Does Not Need to Be Fetched Repeatedly
Another way to reduce rate-limit pressure is to avoid repeatedly requesting slow-changing reference data.
Depending on the integration, examples may include:
- configuration;
- stable metadata;
- mappings;
- definitions that do not change on every execution.
The cache duration should follow the actual freshness requirement of the data.
There is no universal HighLevel cache TTL that fits every resource.
The principle is simpler:
The easiest request to keep under the API limit is the request you never needed to make.
How Should You Retry a GoHighLevel 429?
For a clear 429, start with the information HighLevel returned.
429
↓
inspect quota state
↓
daily exhausted?
├── yes → defer
└── no
↓
burst exhaustion
↓
wait / pace
↓
bounded retryUse bounded retries
Do not retry forever.
A retry policy needs a stopping condition.
Otherwise a temporarily throttled Location can:
- occupy workers indefinitely;
- build an ever-growing backlog;
- increase latency;
- hide a daily-capacity problem;
- turn a local quota issue into a broader system outage.
When retries are exhausted, retain enough information to defer, investigate, or resume the job safely.
Avoid synchronized retry storms
Suppose 100 requests are rejected together.
A naive strategy does this:
100 failures
↓
all wait exactly the same delay
↓
all retry together
↓
another burstRandomized jitter is useful because it spreads retries instead of synchronizing them.
HighLevel's current official JavaScript/TypeScript SDK uses the reported rate-limit interval when available for burst throttling and falls back to exponential backoff with jitter when rate-limit headers are missing. Its retry behavior is opt-in rather than automatically enabled for every application.
That is a useful model even if your integration does not use the official SDK:
Use the API's current rate-limit state first. Use a bounded backoff strategy when you lack better timing information.
Rate-Limit Retry Is Not the Same as Safe Business Retry
This distinction is critical.
A clear API response:
429 Too Many Requestsgives you information about what happened.
A network timeout can be much more ambiguous.
Consider:
Client sends POST
↓
HighLevel processes request
↓
network fails before client receives responseFrom the client's perspective, the operation failed.
But the remote side effect may already have happened.
That is different from a known 429 response.
Therefore:
Do not apply the same retry policy to a clear rate-limit response and to an API write whose outcome is unknown.
For write operations that:
- create records;
- send messages;
- book appointments;
- provision something;
- or produce another external side effect,
confirm whether repeating the endpoint is safe.
Depending on the operation, safe recovery may require an idempotency mechanism, stable external identifier, existence check, or reconciliation with current remote state.
The full unknown-outcome and idempotency architecture belongs in the production API integration guide.
Rate limiting is only one failure mode in a production integration. See the full GoHighLevel API integration guide
Classify the Error Before You Retry It
Not every failed HighLevel request belongs in a retry queue.
Response
Typical first action
400 / 422
Fix request or payload
401
Inspect authentication response; refresh only when expiry is indicated
403
Fix permissions or scopes
404
Fix endpoint/resource/context
429 burst
Pace and use bounded retry
429 daily exhaustion
Defer work
5xx
Potential transient retry
Timeout / no response
Treat the business outcome as potentially unknown
HighLevel's current OAuth FAQ recommends refreshing the access token when the API response indicates that the token has expired, saving the new access and refresh tokens, and then repeating the request.
That means:
Do not assume every 401 means “refresh the token.”Classify the actual authentication error first.
Likewise, repeatedly retrying a malformed 422 request does not make it valid.
Retries should be reserved for failures where another attempt has a plausible reason to succeed.
Monitor Rate Limits Before They Become an Outage
A production integration should make rate-limit behavior observable.
Useful request-level logging can include:
operation
location/resource context
HTTP status
attempt number
X-RateLimit-Max
X-RateLimit-Remaining
X-RateLimit-Daily-Remaining
duration
final outcomeDo not log OAuth access tokens, PITs, API secrets, or unnecessary customer PII.
Operational metrics can also include:
429count;- retry count;
- retry success rate;
- API requests per Location;
- remaining daily capacity;
- queue depth;
- age of the oldest queued job;
- jobs deferred because quota was exhausted.
The useful alert is not always:
“A 429 happened.”
One occasional burst 429 that is correctly paced and recovered may be less important than:
daily remaining declining unexpectedly fastor:
oldest queued job age increasing continuouslyThose signals can reveal capacity problems before the integration stops completing business work.
A Reference Architecture for High-Volume HighLevel API Work
A higher-volume integration may eventually look like this:
Events / jobs
↓
Queue
↓
Identify Location / Company
↓
Resource-specific pacing
↓
Bounded workers
↓
HighLevel API
↓
Read X-RateLimit state
↓
┌───────────────┬───────────────┐
Success Retry later DeferThis is a reference architecture, not a minimum requirement.
A small private integration might only need:
Application
↓
small rate limiter
↓
HighLevel APIThe right architecture depends on:
- workload size;
- number of Locations;
- number of concurrent workers;
- daily request volume;
- acceptable processing delay;
- failure/recovery requirements.
Do not add infrastructure simply because the API has a rate limit.
Add it when the workload makes coordination necessary.
Do You Need Temporal for HighLevel Rate Limits?
No.
A HighLevel 429 is not, by itself, a reason to introduce Temporal.
A queue, scheduler, limiter, or normal worker architecture can often handle:
burst limit
→ wait / pace / retryand:
daily capacity exhausted
→ defer work
→ resume laterDurable orchestration becomes relevant only when the rate-limited HighLevel call is one step inside a much larger long-running business process whose state and recovery must survive extended waits, callbacks, partial failures, and restarts across several systems.
For example:
Start business process
↓
System A
↓
HighLevel call delayed
↓
wait hours/days
↓
external callback
↓
human approval
↓
resume exact process stateIn that situation, rate limiting is only one failure mode inside a broader process-state problem.
The rule remains:
Rate limiting alone does not justify durable orchestration. Durable business-process state might.
GoHighLevel API Rate Limit Production Checklist
Before running a HighLevel API integration at meaningful volume:
- Confirm which API surface and authentication context you are using.
- Do not assume
100 requests / 10 secondsapplies universally. - Read the current
X-RateLimit-*headers returned by the API. - Identify which Location or Company owns the request budget.
- Distinguish burst exhaustion from daily exhaustion.
- Control both request rate and concurrency.
- Coordinate processes that share the same app/resource quota.
- Keep multi-location traffic isolated by resource context.
- Pace requests before repeated
429s become normal behavior. - Bound retry attempts.
- Avoid synchronized retry storms.
- Stop immediate retries when the daily budget is exhausted.
- Reduce unnecessary API calls before only making them slower.
- Prefer events over wasteful polling when the use case supports it.
- Cache slow-changing data where freshness requirements allow.
- Do not treat an ambiguous network timeout like a clear
429. - Confirm that write operations are safe to repeat before replaying them.
- Monitor quota state, retries, queue depth and processing delay.
- Never log credentials or unnecessary customer data.
Need Help With a GoHighLevel API Integration Hitting Rate Limits?
When a production integration starts returning 429s, adding a longer sleep() is not always the real fix.
The underlying issue may be:
- excessive concurrency;
- several workers sharing one Location quota;
- inefficient polling;
- unnecessary API calls;
- a bulk job larger than the daily capacity;
- incorrect retry behavior;
- or a multi-location architecture that does not isolate workload correctly.
Hamza can review the actual workload and determine the smallest architecture that safely handles it.
That may mean adjusting batching or concurrency.
It may require per-Location pacing or queueing.
It may involve reducing the number of API calls.
And when failures go beyond rate limiting, the broader integration may need safer idempotency, reconciliation, or recovery logic.
For rate-limit and API reliability problems, work with a GoHighLevel developer for API integrations
