← All posts

GoHighLevel API Rate Limits: 429 Errors, Retries & Concurrency

Hamza Lamhidra
GoHighLevel API rate limits, 429 errors, throttling, and retries

Understand GoHighLevel API rate limits, X-RateLimit headers, 429 errors, burst vs daily quotas, concurrency control, backoff, and safe retry patterns.

For HighLevel's public API 2.0 endpoints authenticated with OAuth, the current published rate limits are:

Those limits are counted per Marketplace app, per resource, where the resource is a specific HighLevel Location or Company. An app installed on Location A and Location B therefore receives an independent API budget for each resource.

But those numbers should not be treated as a universal limit for every HighLevel API surface.

For example, HighLevel currently documents reduced limits for Sandbox Private Integration Tokens and a separate rate limit for the Agent Studio Public API.

That makes the real production question bigger than:

How many HighLevel API requests can I make?

A reliable integration also needs to know:

The most important rule is simple:

Use the rate-limit state HighLevel returns instead of designing the integration around one hard-coded number.

What Are the Current GoHighLevel API Rate Limits?

For the documented public API 2.0 OAuth context, HighLevel currently publishes the following limits:

Limit

Allowance

Burst

100 requests per 10 seconds

Daily

200,000 requests per day

Scope

Per Marketplace app per Location or Company

HighLevel explicitly scopes those limits by app + resource.

For example:

Marketplace App
      │
      ├── Location A
      │      ↓
      │   its own API budget
      │
      └── Location B
             ↓
          separate API budget

If the application uses its full burst allowance against Location A, that does not mean Location B has consumed the same quota.

But several processes calling the same app + Location A are still operating against the same resource budget.

Does 100 requests per 10 seconds mean exactly 10 requests per second?

No.

Mathematically, 100 requests over 10 seconds has an average of 10 requests per second, but HighLevel documents this as a burst limit over an interval.

The documentation does not tell us to model the system as exactly 10 requests in every individual second, nor does it require developers to reverse-engineer the internal rate-limiting algorithm.

Instead, HighLevel exposes the current quota state in API response headers and says those headers should be treated as authoritative.

So avoid turning:

100 requests / 10 seconds

into an undocumented rule such as:

Exactly 10 requests every second

Your integration should react to the API state HighLevel actually returns.

Which Rate Limit Applies to My HighLevel Integration?

Before building any retry or throttling logic, identify the API surface and authentication context you are actually using.

The common mistake is to see 100 requests / 10 seconds in one HighLevel document and apply it everywhere.

Current official documentation already shows why that is unsafe.

API context

Current documented behavior

Public API 2.0 with OAuth

100 requests / 10 sec + 200,000/day

Sandbox PIT

25 requests / 10 sec + 10,000/day

Production PIT

Standard production limits; verify the applicable production context

Agent Studio Public API

300 requests/minute per sub-account

Other specialized API surfaces

Check that API's current documentation

HighLevel's Sandbox PIT documentation says Sandbox Private Integration Tokens are limited to 25 requests per 10 seconds and 10,000 requests per day. It also says creating multiple PITs does not multiply the sandbox allowance.

HighLevel separately documents 300 requests per minute per sub-account across its Agent Studio Public API endpoints. (HighLevel Support Portal)

Therefore:

“The GoHighLevel API limit is always 100 requests every 10 seconds” is too broad.

A better process is:

Which API am I calling?
        ↓
Which authentication/environment?
        ↓
What does its current documentation say?
        ↓
What quota headers did the actual response return?

That is especially important for integrations that use several HighLevel products or move from Sandbox into production.

Read the X-RateLimit Headers Instead of Counting Requests Yourself

For the documented public API context, HighLevel returns rate-limit information on API responses.

The current documented headers are:

Header

Meaning

X-RateLimit-Limit-Daily

Daily request allowance

X-RateLimit-Daily-Remaining

Requests remaining today

X-RateLimit-Interval-Milliseconds

Length of the burst interval

X-RateLimit-Max

Maximum requests allowed in that interval

X-RateLimit-Remaining

Requests remaining in the current interval

HighLevel specifically recommends reading these values instead of counting requests independently because the response headers are authoritative.

Conceptually:

HighLevel API response
        ↓
X-RateLimit-Max
X-RateLimit-Remaining
X-RateLimit-Daily-Remaining
X-RateLimit-Interval-Milliseconds
        ↓
Update pacing / scheduling decision

This matters because local counters can become inaccurate when:

A production integration should not assume an old blog post knows more about its current quota than the API response itself.

A GoHighLevel API v3 response showing the x-ratelimit headers: daily limit, daily remaining, daily reset, interval, max, and remaining.
A live GoHighLevel API v2 response (public API, OAuth context) returning the current x-ratelimit headers. These values are the authoritative quota state to read, rather than counting requests yourself. Values differ by API surface, so treat these as the OAuth-context numbers, not a universal limit.

A 429 Can Represent Two Very Different Problems

429 Too Many Requests tells you that the current request exceeded an applicable rate limit.

But the correct recovery depends on which limit was exhausted.

Case 1: The burst allowance is exhausted

Imagine a bulk process sends a large number of calls against the same Location in a short period:

500 jobs
   ↓
many workers
   ↓
same HighLevel Location
   ↓
burst budget exhausted
   ↓
429

If daily capacity is still available, this is primarily a pacing problem.

The work may be able to continue after the burst interval allows additional requests.

Case 2: The daily allowance is exhausted

Now consider:

X-RateLimit-Daily-Remaining = 0

Waiting one second, five seconds, or another burst interval does not restore daily capacity.

Continuing to immediately retry only creates more failed work.

The integration instead needs to:

stop immediate retries
        ↓
defer remaining jobs
        ↓
retain their state
        ↓
resume when capacity is available

HighLevel's current official JavaScript/TypeScript SDK demonstrates exactly this conceptual distinction. When its optional rate-limit retry feature is enabled, a burst 429 can wait using the reported interval and retry, while a response showing zero daily requests remaining is surfaced immediately because another immediate attempt cannot succeed.

A useful decision tree is:

429 Too Many Requests
        ↓
Inspect current rate-limit state
        ↓
Daily remaining = 0?
      /           \
    Yes            No
     ↓              ↓
Stop immediate   Burst pressure
retry                 ↓
     ↓           wait / pace
Defer work            ↓
                  bounded retry

Treating every 429 as:

“sleep for a few seconds and try again”

misses this distinction.

Do Not Assume Every HighLevel 429 Includes Retry-After

Many generic HTTP retry examples tell developers to read the Retry-After header.

That is useful when an API provides it, but HighLevel's main Rate Limits documentation does not define Retry-After as the universal mechanism for API 2.0 rate-limit handling.

Instead, it documents the X-RateLimit-* headers.

The current official HighLevel SDK follows the same pattern.

For a burst 429, it uses the reported X-RateLimit-Interval-Milliseconds value. If rate-limit headers are unavailable, the SDK falls back to exponential backoff with jitter.

Therefore:

Respect Retry-After when the specific API you are calling returns and documents it, but do not assume every HighLevel 429 provides that header.

Use the quota information the actual response gives you.

Prevent 429 Errors Instead of Only Reacting to Them

A production integration should not rely exclusively on this pattern:

send as quickly as possible
        ↓
hit 429
        ↓
wait
        ↓
retry

That is reactive throttling.

A better high-volume architecture controls the rate before the API needs to reject traffic.

For example:

Incoming jobs
      ↓
Queue
      ↓
Rate-aware scheduling
      ↓
Bounded workers
      ↓
HighLevel API

This can reduce:

A small integration does not automatically need Redis, distributed workers, or a large queueing platform.

For a workflow making only a few calls, a simple rate limiter may be enough.

The architecture should match the workload.

The principle is:

Pace the work before concurrency turns a normal job into an API burst.

Concurrency and Request Rate Are Not the Same Thing

Two related controls are often confused.

Concurrency

How many API operations are executing at the same time?

Rate

How many requests begin during a given time window?

Those are not identical.

For example, 20 workers do not always produce exactly 20 requests per second. API calls take different amounts of time.

But high concurrency can produce a large burst when requests complete quickly.

Consider:

500 contact updates
        ↓
start everything at once
        ↓
many requests compressed
into a short period

A more controlled model is:

500 jobs
   ↓
bounded concurrency
   +
rate pacing
   ↓
HighLevel API

This matters with:

The goal is not to eliminate concurrency.

It is to keep concurrency from violating the API budget that applies to the resource being processed.

Multiple Workers Calling the Same Location Share the Same Quota

HighLevel scopes its documented OAuth quota by Marketplace app + resource.

It does not describe a separate quota for every application server, container, worker, or process.

That creates an important production consideration.

Suppose three workers all process the same HighLevel Location:

Worker 1 ─┐
Worker 2 ─┼──→ App X + Location A
Worker 3 ─┘

From HighLevel's perspective, that is still traffic against the same app/resource budget.

If each worker independently assumes it owns the entire allowance, the combined traffic can exceed the actual quota.

For larger systems, rate control may therefore need to be coordinated across workers.

Conceptually:

Worker 1 ─┐
Worker 2 ─┼→ shared scheduling / quota coordination
Worker 3 ─┘
                         ↓
                  App X + Location A
                         ↓
                    HighLevel API

That coordination might be implemented with a central queue, shared limiter, scheduler, or another mechanism appropriate to the application.

The implementation is less important than the principle:

Do not give every worker the illusion that it owns a quota HighLevel actually scopes to their shared resource.

Multi-Location Apps Need Per-Location Rate-Limit Thinking

The same scoping rule also creates an advantage for multi-location applications.

For the documented OAuth context:

Marketplace App
      │
      ├── Location A → quota A
      │
      ├── Location B → quota B
      │
      └── Location C → quota C

HighLevel explicitly documents that installing the same app on multiple sub-accounts gives each installation an independent budget rather than dividing one global allowance among all of them.

That means a noisy Location should not automatically force every other Location to stop processing.

A useful scheduling model is:

Incoming jobs
      ↓
identify locationId
      ↓
resource-specific queue/pacing
      ↓
HighLevel

If Location A is temporarily throttled:

Location A
→ slow / defer

while:

Location B
→ continue normally

provided your own infrastructure does not unnecessarily couple them.

This is another reason tenant context needs to remain explicit throughout a multi-location integration.

Daily API Capacity Is a Budget Problem, Not a Backoff Problem

Burst throttling asks:

How quickly can this work be sent?

Daily capacity asks:

Should this process make this many calls at all?

Imagine a job that requires:

5,000 records
×
50 API requests per record
=
250,000 API calls

In the documented OAuth context, that exceeds a 200,000-request daily allowance for one app/resource even if the requests are paced perfectly. (App Marketplace)

No clever backoff algorithm changes the total number of requests.

The useful question becomes:

Why are we making 50 calls for each record?

Common opportunities to reduce usage include:

Pattern

Better question

Polling repeatedly

Can an event/webhook tell us when something changes?

Re-reading stable data

Can the result be cached?

Same lookup for every row

Can the value be reused?

Full sync every run

Can only changed records be processed?

One huge migration

Can the workload be scheduled across safe windows?

Avoid assuming a bulk endpoint exists unless the specific HighLevel API resource documents one.

The goal is not merely to send an inefficient integration more slowly.

It is to reduce unnecessary API work.

Use Webhooks When They Can Replace Wasteful Polling

Polling can consume API budget even when nothing has changed.

For example:

Every minute
    ↓
Fetch records
    ↓
Compare state
    ↓
Nothing changed

Repeated across many Locations, this can become expensive in request count.

When HighLevel exposes an appropriate event, an event-driven design may be more efficient:

HighLevel event
      ↓
Webhook
      ↓
Process the actual change

This does not mean every API read should become a webhook.

It means polling should be justified when an appropriate event-driven boundary exists.

Cache Data That Does Not Need to Be Fetched Repeatedly

Another way to reduce rate-limit pressure is to avoid repeatedly requesting slow-changing reference data.

Depending on the integration, examples may include:

The cache duration should follow the actual freshness requirement of the data.

There is no universal HighLevel cache TTL that fits every resource.

The principle is simpler:

The easiest request to keep under the API limit is the request you never needed to make.

How Should You Retry a GoHighLevel 429?

For a clear 429, start with the information HighLevel returned.

429
 ↓
inspect quota state
 ↓
daily exhausted?
  ├── yes → defer
  └── no
        ↓
     burst exhaustion
        ↓
     wait / pace
        ↓
     bounded retry

Use bounded retries

Do not retry forever.

A retry policy needs a stopping condition.

Otherwise a temporarily throttled Location can:

When retries are exhausted, retain enough information to defer, investigate, or resume the job safely.

Avoid synchronized retry storms

Suppose 100 requests are rejected together.

A naive strategy does this:

100 failures
      ↓
all wait exactly the same delay
      ↓
all retry together
      ↓
another burst

Randomized jitter is useful because it spreads retries instead of synchronizing them.

HighLevel's current official JavaScript/TypeScript SDK uses the reported rate-limit interval when available for burst throttling and falls back to exponential backoff with jitter when rate-limit headers are missing. Its retry behavior is opt-in rather than automatically enabled for every application.

That is a useful model even if your integration does not use the official SDK:

Use the API's current rate-limit state first. Use a bounded backoff strategy when you lack better timing information.

Rate-Limit Retry Is Not the Same as Safe Business Retry

This distinction is critical.

A clear API response:

429 Too Many Requests

gives you information about what happened.

A network timeout can be much more ambiguous.

Consider:

Client sends POST
      ↓
HighLevel processes request
      ↓
network fails before client receives response

From the client's perspective, the operation failed.

But the remote side effect may already have happened.

That is different from a known 429 response.

Therefore:

Do not apply the same retry policy to a clear rate-limit response and to an API write whose outcome is unknown.

For write operations that:

confirm whether repeating the endpoint is safe.

Depending on the operation, safe recovery may require an idempotency mechanism, stable external identifier, existence check, or reconciliation with current remote state.

The full unknown-outcome and idempotency architecture belongs in the production API integration guide.

Rate limiting is only one failure mode in a production integration. See the full GoHighLevel API integration guide

Classify the Error Before You Retry It

Not every failed HighLevel request belongs in a retry queue.

Response

Typical first action

400 / 422

Fix request or payload

401

Inspect authentication response; refresh only when expiry is indicated

403

Fix permissions or scopes

404

Fix endpoint/resource/context

429 burst

Pace and use bounded retry

429 daily exhaustion

Defer work

5xx

Potential transient retry

Timeout / no response

Treat the business outcome as potentially unknown

HighLevel's current OAuth FAQ recommends refreshing the access token when the API response indicates that the token has expired, saving the new access and refresh tokens, and then repeating the request.

That means:

Do not assume every 401 means “refresh the token.”

Classify the actual authentication error first.

Likewise, repeatedly retrying a malformed 422 request does not make it valid.

Retries should be reserved for failures where another attempt has a plausible reason to succeed.

Monitor Rate Limits Before They Become an Outage

A production integration should make rate-limit behavior observable.

Useful request-level logging can include:

operation
location/resource context
HTTP status
attempt number
X-RateLimit-Max
X-RateLimit-Remaining
X-RateLimit-Daily-Remaining
duration
final outcome

Do not log OAuth access tokens, PITs, API secrets, or unnecessary customer PII.

Operational metrics can also include:

The useful alert is not always:

“A 429 happened.”

One occasional burst 429 that is correctly paced and recovered may be less important than:

daily remaining declining unexpectedly fast

or:

oldest queued job age increasing continuously

Those signals can reveal capacity problems before the integration stops completing business work.

A Reference Architecture for High-Volume HighLevel API Work

A higher-volume integration may eventually look like this:

Events / jobs
      ↓
Queue
      ↓
Identify Location / Company
      ↓
Resource-specific pacing
      ↓
Bounded workers
      ↓
HighLevel API
      ↓
Read X-RateLimit state
      ↓
 ┌───────────────┬───────────────┐
Success       Retry later      Defer

This is a reference architecture, not a minimum requirement.

A small private integration might only need:

Application
    ↓
small rate limiter
    ↓
HighLevel API

The right architecture depends on:

Do not add infrastructure simply because the API has a rate limit.

Add it when the workload makes coordination necessary.

Event-driven architecture: a GoHighLevel webhook starts one Temporal workflow per lead, which waits on durable timers, then runs activities on a single worker paced by a per-location token bucket before calling the HighLevel API.
The real pattern is event-driven, not batch. Each HighLevel event starts its own Temporal workflow with a deterministic ID, so duplicate deliveries drop. Work builds up as workflows sitting on durable timers rather than as bulk jobs. When a timer fires, an activity runs on a single worker, and every HighLevel call passes through an in-process token bucket for that location. Temporal is the queue, scheduler, and retry engine.

Do You Need Temporal for HighLevel Rate Limits?

No.

A HighLevel 429 is not, by itself, a reason to introduce Temporal.

A queue, scheduler, limiter, or normal worker architecture can often handle:

burst limit
→ wait / pace / retry

and:

daily capacity exhausted
→ defer work
→ resume later

Durable orchestration becomes relevant only when the rate-limited HighLevel call is one step inside a much larger long-running business process whose state and recovery must survive extended waits, callbacks, partial failures, and restarts across several systems.

For example:

Start business process
      ↓
System A
      ↓
HighLevel call delayed
      ↓
wait hours/days
      ↓
external callback
      ↓
human approval
      ↓
resume exact process state

In that situation, rate limiting is only one failure mode inside a broader process-state problem.

The rule remains:

Rate limiting alone does not justify durable orchestration. Durable business-process state might.

GoHighLevel API Rate Limit Production Checklist

Before running a HighLevel API integration at meaningful volume:

Need Help With a GoHighLevel API Integration Hitting Rate Limits?

When a production integration starts returning 429s, adding a longer sleep() is not always the real fix.

The underlying issue may be:

Hamza can review the actual workload and determine the smallest architecture that safely handles it.

That may mean adjusting batching or concurrency.

It may require per-Location pacing or queueing.

It may involve reducing the number of API calls.

And when failures go beyond rate limiting, the broader integration may need safer idempotency, reconciliation, or recovery logic.

For rate-limit and API reliability problems, work with a GoHighLevel developer for API integrations

Work with Me on Upwork

Book a free workflow audit

30 minutes. We look at your two or three most critical workflows and I tell you which ones are actually at risk. No pitch deck.