API Performance Testing: Strategies for REST and GraphQL Endpoints
API performance testing differs from web page load testing in several fundamental ways. APIs serve machine clients (mobile apps, SPAs, microservices) rather than browsers, which changes the load pattern: no CSS parsing, no image loading, no rendering — just raw request-response cycles. An API endpoint that returns in 50ms may still be a performance problem if it's called 200 times per page load from a frontend application making parallel requests.
The challenge of API performance testing is that APIs vary enormously in their resource cost. A simple key-value lookup endpoint and a complex search endpoint with pagination, filtering, and joins across multiple tables are both "API calls" but may differ in server-side cost by three orders of magnitude. Testing them with the same approach yields misleading results.
Endpoint Profiling
Before writing performance tests, profile your API endpoints to understand their resource costs and usage patterns. Not all endpoints deserve equal testing attention.
Focus your performance testing effort on the top-right quadrant: high-traffic, high-cost endpoints. These are the endpoints where performance regressions have the largest user impact and where capacity limits are most likely to be reached. Common examples include search, feed generation, and checkout flows.
REST API Testing Patterns
CRUD Operation Mix
Real API traffic is a mix of read and write operations, with reads typically dominating. A social media API might see 95% GET requests (reading feeds, profiles, comments) and 5% POST/PUT/DELETE (creating posts, updating profiles). Your test should reflect this ratio because reads and writes have different resource characteristics — reads can be cached and parallelized; writes require serialization and consistency.
// k6 test with realistic CRUD distribution
import http from 'k6/http';
import { check, sleep } from 'k6';
const BASE_URL = 'https://api.staging.example.com/v1';
export default function() {
const roll = Math.random();
if (roll < 0.50) {
// 50% - List with pagination
const res = http.get(`${BASE_URL}/products?page=1&limit=20`, {
tags: { operation: 'list' }
});
check(res, {
'list status 200': (r) => r.status === 200,
'list has items': (r) => JSON.parse(r.body).data.length > 0,
});
} else if (roll < 0.80) {
// 30% - Get single item
const id = Math.floor(Math.random() * 10000) + 1;
const res = http.get(`${BASE_URL}/products/${id}`, {
tags: { operation: 'get' }
});
check(res, {
'get status 200': (r) => r.status === 200,
});
} else if (roll < 0.90) {
// 10% - Search (most expensive)
const terms = ['laptop', 'phone', 'headphones', 'tablet', 'charger'];
const q = terms[Math.floor(Math.random() * terms.length)];
const res = http.get(`${BASE_URL}/products/search?q=${q}&limit=20`, {
tags: { operation: 'search' }
});
check(res, {
'search status 200': (r) => r.status === 200,
});
} else if (roll < 0.97) {
// 7% - Create
const res = http.post(`${BASE_URL}/products`,
JSON.stringify({ name: `Product ${Date.now()}`, price: 29.99 }),
{ headers: { 'Content-Type': 'application/json' },
tags: { operation: 'create' }
}
);
check(res, {
'create status 201': (r) => r.status === 201,
});
} else {
// 3% - Update
const id = Math.floor(Math.random() * 10000) + 1;
const res = http.put(`${BASE_URL}/products/${id}`,
JSON.stringify({ price: 39.99 }),
{ headers: { 'Content-Type': 'application/json' },
tags: { operation: 'update' }
}
);
check(res, {
'update status 200': (r) => r.status === 200,
});
}
sleep(Math.random() * 2 + 0.5); // 0.5-2.5s think time
}
Pagination Testing
Pagination performance degrades with offset depth. Page 1 is fast; page 500 is often slow because OFFSET 10000 still requires the database to scan and skip 10,000 rows. Test deep pagination explicitly to ensure that users (or crawlers) browsing deep into lists do not hit query timeouts. Cursor-based pagination avoids this problem by using an indexed column as the cursor instead of a row offset.
Rate Limiting Validation
API rate limits protect your backend from abuse and ensure fair access. Performance tests should validate that rate limiting works correctly under load — that rate-limited responses (HTTP 429) are returned quickly (they should not consume significant server resources) and that non-rate-limited requests continue to perform normally even while some clients are being throttled.
GraphQL Performance Considerations
GraphQL introduces performance challenges that REST APIs do not have. The flexible query language means clients can construct queries of arbitrary complexity — from a simple field lookup to a deeply nested query that joins dozens of tables and returns megabytes of data.
Query Complexity
Test with the full range of query complexity that your schema allows. If your schema permits a query like users { posts { comments { author { posts } } } }, that deeply nested query is a valid performance test case even if no current client sends it. A future client (or a malicious actor) will.
// k6 GraphQL performance test with varying complexity
import http from 'k6/http';
const QUERIES = {
simple: `query { user(id: 1) { name email } }`,
medium: `query {
users(limit: 20) {
name email
posts(limit: 5) { title createdAt }
}
}`,
complex: `query {
users(limit: 50) {
name email avatar
posts(limit: 10) {
title body
comments(limit: 5) {
text author { name }
}
}
}
}`,
};
export default function() {
const roll = Math.random();
let query;
if (roll < 0.60) query = QUERIES.simple;
else if (roll < 0.90) query = QUERIES.medium;
else query = QUERIES.complex;
http.post('https://api.staging.example.com/graphql',
JSON.stringify({ query }),
{ headers: { 'Content-Type': 'application/json' } }
);
}
N+1 Query Detection
GraphQL resolvers are prone to N+1 query problems. A query for 20 users with their posts triggers 1 query for users + 20 queries for posts (one per user). DataLoader batching solves this, but only if it is correctly implemented. Performance tests that request lists with nested relations are the most effective way to detect N+1 regressions — the response time will scale linearly with list size instead of staying constant.
Authentication and Session Handling
API performance tests must handle authentication realistically. An unauthenticated test bypasses token validation, session lookup, and authorization middleware — which can be a significant fraction of total request processing time. Use a pool of pre-created test accounts with valid tokens, and rotate through them to simulate realistic per-user isolation.
// Token pool for authenticated API testing
import http from 'k6/http';
import { SharedArray } from 'k6/data';
const tokens = new SharedArray('tokens', function() {
return JSON.parse(open('./test-tokens.json'));
});
export default function() {
const token = tokens[__VU % tokens.length]; // Distribute across VUs
http.get('https://api.staging.example.com/v1/me', {
headers: {
'Authorization': `Bearer ${token.access_token}`,
'Content-Type': 'application/json',
},
});
}
Payload Size and Serialization
Large response payloads stress both the server (serialization time, memory allocation) and the network (transfer time, bandwidth). Test your API with realistic response sizes:
| Payload Size | Serialization Impact | Network Impact (3G) | Testing Focus |
|---|---|---|---|
| < 10 KB | Negligible | < 100ms | Standard latency testing |
| 10-100 KB | Moderate (JSON.stringify) | 100ms-1s | Serialization efficiency, compression |
| 100 KB-1 MB | Significant (memory allocation) | 1-10s | Pagination vs. bulk, streaming |
| > 1 MB | High (GC pressure) | > 10s | Should this be a download endpoint? |
Enable compression (Accept-Encoding: gzip) in your performance tests and verify that the server returns compressed responses. JSON compresses well (typically 5-10x), and the CPU cost of gzip compression is almost always less than the network transfer time savings, especially for mobile clients.
Error Handling Under Load
APIs should fail gracefully under load. Error responses (4xx, 5xx) should be returned quickly — they should not consume the same resources as successful responses. Common failure: a slow database query causes a timeout, but the server holds the connection open for the full timeout duration (30 seconds) before returning 504. Under load, these held connections accumulate and consume the entire connection pool.
Monitoring API Performance in Production
Performance tests establish baselines and catch regressions. Production observability validates that test results predict real-world behavior. Track these metrics per endpoint in production:
- Response time percentiles (p50, p95, p99): Per endpoint, per method. A global average across all endpoints is meaningless.
- Error rate by type: Distinguish between client errors (4xx — may indicate API contract issues) and server errors (5xx — indicate backend failures).
- Throughput: Requests per second per endpoint. Sudden changes indicate either traffic shifts or upstream client changes.
- Payload sizes: Track response body size distributions. A sudden increase may indicate a query that returns more data than expected (missing pagination, removed filters).
Key Takeaways
- Profile endpoints by traffic volume and resource cost. Focus performance testing on high-traffic, high-cost endpoints (search, feeds, checkout) — the "critical path" quadrant.
- Model realistic CRUD distributions: reads typically dominate (90-95%), and the read/write ratio affects caching behavior and database contention patterns.
- Test deep pagination explicitly — offset-based pagination degrades at depth, and cursor-based pagination should be validated to not have the same problem.
- For GraphQL, test with varying query complexity and watch for N+1 query regressions by measuring response time scaling with list sizes.
- Authenticate performance tests with real token pools. Unauthenticated tests bypass token validation, session lookup, and authorization — a significant fraction of real processing time.
- Error responses should be fast. A slow error (30-second timeout held open) is worse under load than a fast error because it consumes connection pool resources.