scriptRun BetascriptRun Beta
  • Marketplace
  • Workflows
  • Projetos
  • API
  • Preços
  • Blog
  • Casos de uso
pt
Marketplace
Páginas legaisContatos
Português (PT)
Terms of ServicePrivacy PolicyCookie Policy
© 2024-2026 PIXEL LORIS IT SERVICES L.L.C. All rights reserved.
  1. Blog
  2. API latency: How to measure, benchmark, and reduce it

API latency: How to measure, benchmark, and reduce it

ArchieArchie
5 outubro 2026
Business
AI WorkflowsAutomation
AI Workflows
Automation
+2
AI latency: Find bottlenecks and reduce response times

Índice:

TL;DR
What is API latency?
Where does the delay happen in an AI API request?
How to measure API latency step by step
What is a good API latency for your application?
How to run a meaningful DeepL API latency benchmark?
What does “Discord elevated API latency” mean?
How to reduce API latency without sacrificing reliability?
How do you keep performance consistent as usage grows?

API latency: How to measure, benchmark, and reduce it

An API can return a successful response and still leave users waiting. A support assistant pauses before answering, a translation takes too long to appear, or a dashboard waits for data. API latency helps explain the difference between a request completing successfully and an application responding quickly enough for its users.

Finding the cause takes more than checking a response code. Delays can come from connection setup, server queues, processing, or external services. Several API calls running in sequence can also extend the total wait. Changing providers without measuring these steps may leave the actual bottleneck untouched.

This guide explains what to measure, how to establish a baseline, and how to investigate slow requests. It also covers ways to reduce delays while checking their effect on reliability and output quality.

TL;DR

  • API latency is the time from starting a request to receiving the complete response. Teams define it differently, so check where a measurement starts and stops before comparing numbers.
  • Measure real requests from the client side. Use curl's timing fields to see DNS, connection, TLS, first-byte, and total time. For streaming AI responses, track time to first text separately from total completion time.
  • Report percentiles, not averages. Use p50, p95, and p99 with the sample size and error rate. With only 100 requests, p99 rests on about one result, so use larger samples for tail latency decisions.
  • Test under realistic load. Send requests at a fixed arrival rate, separate fresh and reused connections, and keep failures and retries in your results.
  • For DeepL benchmarks, always set source_lang, set model_type explicitly, and record model_type_used, or you may end up comparing the same model twice.
  • For Discord delays, check the status page and your own logs, respect rate-limit headers, and acknowledge slash commands within three seconds.
  • To reduce API latency, reuse connections, cut unnecessary processing, trim AI prompts and output limits, and control retries with backoff and jitter. Shorter prompts and outputs also support LLM API cost optimization. Change one thing at a time and confirm errors and output quality haven't gotten worse.

What is API latency?

API latency is the delay between initiating an API request and receiving its response. In this guide, we measure it from the start of the client request until the complete response arrives, usually in milliseconds.

Terminology varies. Some teams use latency and response time interchangeably; others reserve latency for network travel and use response time for the complete exchange. Always check where a measurement starts and stops before comparing numbers.

Metric

What it measures

Network latency

Travel delay across a network; specify one-way or round-trip

Time to first byte (TTFB)

Time from request initiation to the first response byte

Total response time

Time until the complete response arrives

Server processing time

Time spent executing work on the server

Throughput

Work completed per unit of time, such as requests per second

For an illustrative request that finishes in 800 milliseconds, server processing might account for 500 milliseconds. Connection setup, transit, queueing, and response delivery account for the remainder.

That distinction matters because faster server code addresses only part of the delay. Likewise, higher throughput does not guarantee that each individual request finishes sooner.

Where does the delay happen in an AI API request?

When your application accesses models through ScriptRun’s OpenAI-compatible API, measure the complete request from your application, including the gateway and the model response. This establishes a baseline before you change models, payloads, or timeout settings.

An AI request can pass through several stages: the client prepares its input, establishes a connection, and sends the request. A gateway may authenticate and route it. The provider then queues and processes the input before returning generated output.

ScriptRun brings multiple model providers behind a unified interface and supports response streaming. The useful performance question is how that complete route behaves with your application’s actual inputs, rather than how quickly a model responds in isolation. 

API latency also differs from the time needed to complete a business task. A support assistant might retrieve customer information, generate an answer, and save a record. If those operations depend on each other, their delays accumulate. Measure each call and the overall task so you can identify the step holding up the result.

How to measure API latency step by step

To measure API latency accurately, time real requests from your application's side and report the results as percentiles rather than a single number. The process has four steps: define the exact request and where timing starts and stops, record client-side timings for each stage, repeat the test under realistic traffic, and report percentiles with the sample size. Each step removes a common source of misleading results, such as comparing different request types or trusting one fast response.

Define the request and the measurement boundary

Start with a request your application actually makes. Record the endpoint, HTTP method, payload size, client region, authentication method, and connection-reuse behavior. For AI requests, also record the model, input length, output limit, and streaming setting.

Use a monotonic clock when timing requests in application code. For streaming text responses, measure from request initiation to the first chunk containing non-empty generated text, separately from total completion time. Call this “time to first text” in your results. Ignore role-only or empty chunks; receiving headers or the first stream event does not necessarily mean generated text has arrived. 

Record timings from the client

The following Bash template measures a JSON POST request. Set API_URL and API_KEY in your environment and create a valid request.json payload first. This example assumes bearer authentication; adapt it to your API.

</> bash

curl --silent --show-error \

  --output /dev/null \

  --connect-timeout 10 \

  --max-time 60 \

  --request POST "$API_URL" \

  --header "Authorization: Bearer $API_KEY" \

  --header "Content-Type: application/json" \

  --data-binary @request.json \

  --write-out 'http_status=%{http_code}\ndns_s=%{time_namelookup}\nconnect_s=%{time_connect}\ntls_s=%{time_appconnect}\nttfb_s=%{time_starttransfer}\ntotal_s=%{time_total}\n'

Use the ScriptRun API documentation for its endpoint, authentication, and supported request structure.

These values are in seconds. DNS, connection, TLS, and first-byte timings are cumulative milestones measured from the start, so do not add them together. The curl timing reference defines each field. TTFB includes more than network travel, and client timings alone cannot isolate provider queueing from model processing.

A zero tls_s value does not automatically indicate a problem. Plain HTTP has no TLS handshake, and a reused connection may not require a new handshake. 

Validate a sample response separately: the command discards the body. Check curl’s exit status as well as the HTTP status. If curl receives no HTTP response code, http_status can show 000; this is not a status returned by the API. A timeout after response headers arrive may still show an HTTP code, so inspect the exit status in every case. The timeout values are examples, not recommended service targets.

Repeat the test under realistic conditions

One request is a spot check. Repeat the workload at normal and expected peak traffic, within your provider’s limits. Test fresh connections and reused connections separately; launching a new curl process each time does not reproduce a persistent application connection pool.

If production requests arrive independently, use a load test that starts requests at a configured arrival rate rather than waiting for each response before starting the next. Otherwise, a slowdown reduces the load generated by the test and can understate delays under sustained demand. Record the achieved arrival rate and any scheduled requests the load generator could not start. Sequential tests are still useful for baselines and workloads that genuinely wait between requests. 

Separate cache hits from misses. Record retries and failures instead of removing them from the results.

Report percentiles and sample size

Report p50 for the typical request, p95 for the slower portion, and p99 for the tail. A p95 of 900 milliseconds means approximately 95% of measured requests finished within that time. P99 is not the maximum.

Include sample size, test duration, arrival rate, concurrency, and error rate. With 100 requests, the upper 1% contains roughly one observation, making p99 sensitive to individual results. Even 1,000 requests provide only about ten observations in that tail. Treat 1,000 requests per condition as an initial sample, not a guarantee of a stable p99; use larger samples and repeated runs for decisions that depend on tail latency. If percentiles cover successful requests only, label them and report failures separately. 

What is a good API latency for your application?

A good target depends on what the user is trying to do. An interactive search needs a result quickly enough to keep the person engaged. A background translation can take longer if the interface clearly shows that work is continuing.

Request type

What the user waits for

What to measure

Interactive lookup

A complete result

P95 total response time and errors

Streaming AI response 

Generated text begins arriving 

Time to first text and total duration 

Background task

Finished work

Queue delay and completion time

Start with the acceptable wait for the complete action, then allocate time to its dependent steps. In a hypothetical two-second interaction, three sequential calls cannot each consume the full two seconds. Leave room for application logic and rendering, too.

Choose targets by endpoint and workload rather than applying one threshold to everything. Compare the same request class over time. Combining short lookups with long document-processing jobs can make an overall average look alarming or reassuring for the wrong reasons.

How to run a meaningful DeepL API latency benchmark?

A useful benchmark starts with a fixed translation workload. Choose representative texts and keep their language pairs, lengths, and formatting consistent between runs. Otherwise, a faster result may simply reflect an easier request.

If you adapt the curl example above for DeepL, replace the bearer header with Authorization: DeepL-Auth-Key $API_KEY. Set the request URL to https://api-free.deepl.com/v2/translate for API Free or https://api.deepl.com/v2/translate for API Pro. Validate a successful translation before collecting benchmark timings. 

Record the following conditions before testing:

  • API endpoint and account tier.

  • Explicit source_lang and target_lang values, plus text length. Without source_lang, DeepL uses its quality-optimized models regardless of model_type, so test automatic language detection as a separate configuration.

  • Requested model_type, returned model_type_used, and any glossary, context, tag-handling, style, or custom-instruction settings. 

  • Client region and connection reuse.

  • Texts per request and concurrent requests.

  • Test date, duration, sample size, and failures.

DeepL’s text translation reference documents latency_optimized and quality_optimized. These settings express a performance or quality preference rather than selecting a fixed model. Different settings may use the same underlying model for some requests. Set model_type explicitly and record model_type_used for each translation in every successful response. Compare the observed speed and quality without assuming the settings always select different models. 

Use a results sheet like this. The empty cells are for measurements from your own test, not published DeepL performance figures.

Test conditions

Requests

P50 (ms)

P95 (ms)

P99 (ms)

Errors/timeouts

Single text, reused connection, concurrency 1

Fill in

Fill in

Fill in

Fill in

Fill in

Same workload, higher concurrency

Fill in

Fill in

Fill in

Fill in

Fill in

Batched texts, recorded batch size

Fill in

Fill in

Fill in

Fill in

Fill in

Keep the endpoint, region, and translation settings attached to every result. For batches, also calculate characters translated per second. A batch can improve throughput while making each response take longer.

Review translation quality alongside timing. Faster output only helps if it still meets the task’s requirements. Repeat runs at different times before drawing conclusions about consistent performance.

What does “Discord elevated API latency” mean?

Check which metric you are looking at. For example, discord.py’s Client.latency measures the WebSocket gateway heartbeat round trip, not the duration of a REST API request. Measure outgoing HTTP calls separately when investigating slow message sends or other REST operations. 

The phrase describes API responses taking longer than expected. It does not identify the cause by itself, and it does not necessarily mean every Discord feature is unavailable.

First, check the Discord service status for an incident matching the time and affected component. Compare that information with your own logs. A provider incident is one possible explanation; application queues and local network problems can produce similar symptoms.

Next, identify which operations are slow. Record their timestamps, response codes, and durations. A slow successful response differs from an HTTP 429 response indicating a rate limit. Your application may also spend time waiting before it sends the request at all.

Follow Discord’s rate-limit guidance to schedule requests using the returned headers, including X-RateLimit-Bucket, X-RateLimit-Remaining, and X-RateLimit-Reset-After. Stop sending requests to an exhausted bucket until it resets. If a request still receives HTTP 429, respect the returned Retry-After header or retry_after field before retrying. 

For a bot, compare the complete interaction time with the duration of each outgoing API call. If the call is quick but the bot responds late, inspect its task queue and other dependencies. Record any retry waiting separately so it remains visible in the user’s total wait.

Slash commands and button interactions also have a response deadline: send an initial response within three seconds of receiving the interaction. If the work will take longer, send a deferred acknowledgment within that window, then edit the original response or send a follow-up when the result is ready. Edits and follow-up messages remain possible for 15 minutes after the interaction. This deadline does not apply to ordinary message handling. 

How to reduce API latency without sacrificing reliability?

To reduce API latency without sacrificing reliability, start with the stage your measurements show is slowest, then change one thing at a time. Most delays fall into four areas: connection and transfer overhead, unnecessary processing, oversized AI requests, and poorly controlled retries or fallbacks. After each change, rerun the same workload and confirm that errors and output quality have not gotten worse, because a faster response only helps if it is still correct.

Reduce connection and transfer overhead

Reuse connections through a suitable client or connection pool. Avoid repeating connection setup for every call when the API and client support persistent connections.

Test deployment regions that make sense for your users and downstream providers. Reduce oversized payloads and request only the fields needed. If a query is dominated by server processing, however, a smaller response alone will have limited effect.

Remove unnecessary processing

Inspect slow database queries, repeated lookups, and duplicate API calls. Cache results when freshness, permissions, and invalidation requirements allow it.

Run independent operations concurrently, with a limit on simultaneous work. Dependencies still need their required inputs, and launching unlimited requests can move the bottleneck into a queue or trigger rate limits.

For multi-step implementations, the ScriptRun workflow builder provides branching, API triggers, webhooks, and execution logs to organize calls and examine how a run progresses.

Adjust AI requests to the task

Choose a model that meets the required quality level. Remove irrelevant prompt material and set an output limit appropriate to the answer. Keep instructions and evidence the model needs to produce a correct result.

Streaming can make generated text available earlier, but it does not guarantee a shorter total completion time. Compare time to first text, final quality, and cost per completed task before accepting a change.

Control queues, retries, and fallbacks

Set timeouts according to the time available for the whole operation. Limit retry attempts and total retry time, and use exponential backoff with random jitter for retryable failures so clients do not all retry together. Honor provider-specified retry delays; jitter must not cause an earlier retry. Retry only when the operation is safe to repeat, using documented idempotency support for writes where available. 

A fallback may rescue a failed request while adding another wait. Measure that path separately from first-attempt successes. To reduce API latency sustainably, make one targeted change and repeat the original workload, checking that errors and incorrect outputs have not increased.

How do you keep performance consistent as usage grows?

Monitor latency by endpoint, model, region, and request type. Track percentiles beside traffic, errors, timeouts, and retry counts. Compare results before and after deployments so a regression has a clear starting point.

Keep test inputs representative of production. A short prompt does not reveal how the application handles a long conversation, and one successful call does not show what happens when requests arrive together.

If you install prebuilt AI workflows, test the complete run with realistic inputs before increasing usage. Individual nodes can perform well while the overall sequence remains too slow.

Start with one representative request in ScriptRun. Record its API latency, output quality, and cost. Change one setting, rerun the same workload, and keep the change only if the complete result improves.


Útil?
Compartilhe este artigo!

Crie workflows de IA em minutos

Crie automações, conecte APIs e acesse mais de 50 modelos de IA em uma única plataforma.

Construtor visual de workflows
API compatível com OpenAI
Marketplace de workflows e templates
Comece grátis, planos a partir de $5

Índice:

TL;DR
What is API latency?
Where does the delay happen in an AI API request?
How to measure API latency step by step
What is a good API latency for your application?
How to run a meaningful DeepL API latency benchmark?
What does “Discord elevated API latency” mean?
How to reduce API latency without sacrificing reliability?
How do you keep performance consistent as usage grows?

Crie workflows de IA em minutos

Crie automações, conecte APIs e acesse mais de 50 modelos de IA em uma única plataforma.

Construtor visual de workflows
API compatível com OpenAI
Marketplace de workflows e templates
Comece grátis, planos a partir de $5

FAQ

Is API latency the same as response time?

Sometimes. Teams use the terms differently. This guide measures from client request initiation to receipt of the complete response. Before comparing dashboards or providers, confirm that their timings use the same start and end points.

Why is p95 more useful than an average alone?

P95 helps show the experience of slower requests that an average can hide. Use it alongside the median, errors, and sample size. It does not describe the slowest request or explain why delays occur.

Can I measure an API using ping?

Ping measures an ICMP network round trip, where allowed. It does not execute your authenticated API request, database query, or model inference. Measure a real application request to understand the delay your users encounter.

Does streaming reduce the total time an AI request takes?

Not necessarily. Streaming lets an application receive text before generation finishes, which can improve responsiveness. Measure time to the first non-empty text chunk and full completion separately. Also check whether the interface displays incoming text promptly.

Why do requests become slower under load?

Requests can wait for available connections, workers, database access, or provider capacity. Once a constrained resource cannot keep up, queues grow. Compare latency with traffic and resource measurements to locate the constraint before adding capacity.

Can retries make an application slower?

Yes. Each additional attempt and waiting period extends the complete operation. Retries can help recover from temporary failures, but they need limits and a total deadline. Report first-attempt timing separately from operations that required retries.

Does adding an API gateway affect performance?

It can. A gateway adds routing and processing to the request path, while its connection handling or caching may change overall behavior. Measure the full route with consistent inputs rather than assuming a fixed performance penalty or benefit.

How should I test performance through ScriptRun?

Use requests that represent your application, with consistent model settings, inputs, and output limits. Measure timing, errors, quality, and cost across repeated runs. For workflows, record both individual call durations and the time to finish the complete task.

Leia mais do nosso blog

Publicações relacionadas

AI MarketplaceAI Workflows
AI Marketplace
AI Workflows
+2
AI Coding Assistants Like Claude Code: 2026 Rankings
Business 8 outubro 2026

AI Coding Assistants Like Claude Code: 8 Ranked for 2026

Compare AI coding assistants like Claude Code in 2026: real pricing, usage limits, and the right choice for terminal, IDE, budget or enterprise teams.

4
TinaTina
AI WorkflowsAutomation
AI Workflows
Automation
+2
Is ChatGPT a General-Purpose Technology? Evidence & Uses
Business 5 outubro 2026

Is ChatGPT General Purpose Technology?

Is ChatGPT General Purpose Technology? Explore the evidence, what it means for businesses, and how ScriptRun helps turn AI capabilities into repeatable processes.

6
ArchieArchie
AI MarketplaceAI Workflows
AI Marketplace
AI Workflows
+2
Best AI Coding Assistant 2026: 8 Tools Ranked by Job
Business 3 outubro 2026

Best AI Coding Assistant in 2026: 8 Tools Compared by Job

Compare 8 AI coding assistants by job, not hype: current 2026 pricing, agent features, privacy options and the limits vendors leave out of their pages.

12
TinaTina
AI WorkflowsAutomationPrompt Engineering
AI Workflows
Automation
Prompt Engineering
+3
LLM API Cost Optimization & Token Pricing Guide | ScriptRun
Business 30 setembro 2026

LLM API Cost Optimization: Strategies and Unified AI Billing

Practical strategies for LLM API cost optimization. Learn how to reduce your cloud computing bills using dynamic model routing, context caching, and a unified enterprise billing platform.

14
James