Skip to main content
These settings size the API and Brainstore compute for your workload. The defaults are suitable for most deployments. Tune them for high ingestion volume, heavy eval workloads, or cluster capacity constraints.

Configure AWS API ECS services

On AWS, Terraform module v6.0 introduces ECS as the target runtime for the API, split into three services by workload type. Each service scales automatically through Application Auto Scaling, adjusting its running task count between a minimum and maximum based on multiple metrics, including CPU target tracking and event loop timing. The module exposes optional variables to tune per-task size and task counts. The defaults are suitable for most deployments. See Upgrade to Terraform module v6 for the migration. 1024 CPU units equal 1 vCPU. braintrust-api handles general API traffic. braintrust-api-ingest handles ingestion paths (/logs3, /logs3/overflow, /otel/v1/*). braintrust-api-background handles background paths (evals, function invocation, proxy). For example, to raise the ingestion service’s floor for a high-volume deployment:
Raise braintrust_api_ingest_min_count for deployments with spiky ingestion volume or if you expect a launch that will suddenly increase log traffic, and braintrust_api_background_min_count for heavy eval workloads.

Configure Helm API workload isolation

This feature is available to Kubernetes deployments using Helm chart 6.13.0+. AWS ECS deployments use the Terraform-managed routing described above.
Setting api.workloadIsolation.enabled: true creates dedicated braintrust-api-ingest and braintrust-api-background Deployments and Services alongside the default braintrust-api pool. The pools share the same image and base configuration while allowing independent replica counts, resources, probes, rollout settings, and disruption budgets, so classified ingestion or background load does not consume the default API pool’s capacity. Your ingress or gateway must keep braintrust-api as its default backend and route the following paths to the specialized pools: If your ingress cannot match requests by HTTP method (for example, GKE Ingress), route the listed paths for all methods instead. Pools use fixed replica counts by default. Configure them with api.replicas, api.workloadIsolation.ingest.replicas, and api.workloadIsolation.background.replicas. On GKE, you can scale each pool automatically instead. See Configure GKE API autoscaling.

Roll out workload isolation for an existing deployment

For an existing deployment, stage the rollout so you can verify each pool before routing traffic to it:
  1. Create the pools without routing traffic to them or switching Brainstore’s internal AI proxy. Set api.workloadIsolation.brainstoreAiProxyToBackground: false and apply:
  2. Verify the ingest and background pools are ready. Then update your ingress or gateway to route the classified paths to the new Services, and set brainstoreAiProxyToBackground: true to switch Brainstore’s internal AI proxy to the background pool:
To roll back, first return the classified paths and brainstoreAiProxyToBackground to the default API Service and verify it is serving them, then disable workload isolation in the chart. The chart cannot update an external ingress or gateway on its own, so disabling isolation before rerouting would leave the classified paths pointed at Services that no longer exist.

Integrate with the Istio VirtualService

When using the chart-managed Istio VirtualService, set virtualService.workloadIsolation.enabled: true after the isolated pools are healthy. The chart then renders the ingest and background route contract before your existing virtualService.http rules. The classified routes take precedence, so do not use virtualService.http to override a classified path.
Enable this only after the ingest and background pools are running and ready.
This option requires both virtualService.enabled: true and api.workloadIsolation.enabled: true. The chart fails to render if either is missing.

Configuration reference

Default pool availability settings

Helm chart 6.13.0 or later also exposes availability settings for the default braintrust-api pool. The defaults preserve the chart’s previous behavior, so no configuration changes are required when upgrading. The ingest and background pools inherit these values as their base configuration.

Configure GKE API autoscaling

This feature is available to GKE deployments using Helm chart 6.15.0 or later. It is built on GKE’s AutoscalingMetric resource, which Google classifies as Preview (Pre-GA), so Google may change or discontinue it.
Setting api.autoscaling.enabled: true renders a HorizontalPodAutoscaler and an AutoscalingMetric for each API pool. Each pool then scales on three signals:
  • CPU utilization scoped to the api container, so sidecars and extraContainers do not skew the measurement.
  • Node.js event-loop utilization, as a ratio from 0 to 1.
  • Mean Node.js event-loop delay, in seconds.
These are the same three signals AWS uses through Application Auto Scaling. Default CPU (50%) and event-loop utilization (40%) match. Event-loop delay is the exception: AWS uses step scaling, which Kubernetes HPA does not support, so GKE uses target tracking at 50 ms instead. See Configure AWS API ECS services.

Prerequisites

  • Data plane v2.9.0 or later, which serves the Prometheus metrics endpoint on the API health server.
  • GKE 1.35.1-gke.1396000 or later.
  • The Performance HPA profile enabled on the cluster.
  • The Autoscaling API enabled on the cluster.
  • roles/autoscaling.metricsWriter granted to every node service account.
  • The Autoscaling API included in your service perimeter, if you use VPC Service Controls.
helm install and helm upgrade fail fast if the chart detects a missing prerequisite, rather than rendering a partial configuration:
  • Setting cloud to anything other than google fails, because autoscaling is supported only on GKE.
  • A cluster without the autoscaling.gke.io/v1beta1 API fails. Verify with kubectl api-resources | grep autoscalingmetric.

Enable autoscaling

When autoscaling is enabled for a pool, the chart omits replicas from that pool’s Deployment and the HorizontalPodAutoscaler controls the replica count instead. The chart also sets ENABLE_PROMETHEUS_METRICS to true and exposes api.healthServer.port (8001 by default) on the api container so the AutoscalingMetric can scrape it. Account for that port in any network policy that restricts traffic to the API pods.

Override autoscaling per pool

With api.workloadIsolation.enabled: true, the ingest and background pools inherit every value under api.autoscaling and can override any of them under api.workloadIsolation.<pool>.autoscaling, including the metric targets and the scaling behavior:
This example turns on workload isolation at the same time as autoscaling, which suits a new deployment. On an existing deployment, stage the pools first. The example leaves api.workloadIsolation.brainstoreAiProxyToBackground at its default true, which switches Brainstore’s internal AI proxy to the background pool as soon as you apply it. See Roll out workload isolation for an existing deployment. Setting api.workloadIsolation.<pool>.autoscaling.enabled: false keeps that pool at a fixed replica count while its siblings continue to scale. That pool’s replicas value is honored again.
The Helm chart ships examples/google-api-isolation-autoscaling/values.yaml as a starting point for this pattern. Combine it with an Autopilot or Standard values file, which supply the rest of the configuration.
For background on the GKE resource behind this feature, see Google’s Expose custom metrics for autoscaling.

Configuration reference

Configure Brainstore fast readers

Fast readers are isolated Brainstore nodes dedicated to serving predictable UI queries (paginated viewers, span and trace lookups), preventing resource-intensive ad-hoc queries from making the UI unresponsive.
  • GCP and Azure: Fast readers are enabled by default starting in Helm chart v5.0.0. See the configuration reference below.
  • AWS: Fast readers are enabled by default (2 nodes) starting in Terraform module v5.5.0. On earlier module versions they are disabled by default. Set brainstore_fast_reader_instance_count in your Terraform configuration to control the node count, or set it to 0 to opt out (recommended for sandbox or non-production deployments).
Upgrading to Helm chart v5.0.0 from an earlier version automatically creates fast reader nodes. By default, 2 fast reader nodes are created with the same resource profile as standard reader nodes (CPU: 16, memory: 32Gi). Verify that your cluster has capacity for these additional nodes before upgrading.
If you have custom brainstore.readinessProbe overrides pointing to /status, remove them before upgrading to Helm chart v5.0.0+. The /status readiness endpoint has a bug where it never recovers after a failure, which can permanently mark Brainstore nodes as not ready. Remove any brainstore.readinessProbe or brainstore.fastreader.readinessProbe customizations and rely on the chart defaults.

Configuration reference

Fast readers are configured under the brainstore.fastreader key in your values.yaml. If you have customized brainstore.reader settings, mirror those customizations to brainstore.fastreader.
Azure users must explicitly set brainstore.fastreader.volume.size when using Azure Container Storage (enableAzureContainerStorageDriver: true):

Brainstore rollout controls

This section applies to GCP and Azure deployments using the Helm chart (6.17.1+). AWS deployments manage Brainstore rollout through Terraform and do not expose these settings.
Brainstore readers, fast readers, and writers each have independently configurable rollout strategy and readiness dwell settings. These controls let you limit how many replacement pods start together and require replacements to stay ready for a minimum duration before a rollout continues. This can reduce simultaneous cache warm-up, object-storage, scheduling, and compaction pressure on cache-heavy or high-throughput deployments. The defaults preserve the rollout behavior from earlier chart versions and apply to all three Brainstore roles: Replace <role> with reader, fastreader, or writer. Upgrading without overriding these values does not change rollout pacing. These settings do not change the pod template and do not start a rollout on their own. Kubernetes uses updated settings for an active rollout and future rollouts. For cache-heavy or high-throughput deployments, use a longer readiness dwell:
progressDeadlineSeconds must be greater than minReadySeconds. The chart rejects an invalid pairing. Keep enough deadline margin for scheduling, image pulls, startup, and the readiness dwell. Longer readiness dwell reduces rollout pressure but does not confirm that a pod’s local cache is fully warm. Monitor workload health until the rollout converges. maxUnavailable: 0 requires enough capacity to schedule the configured surge. On managed clusters, even maxSurge: 1 can require an additional node and local storage. Before upgrading, verify regional compute quota and the availability of the required node type and storage, and account for the temporary increase in compute and storage cost.
Setting strategy.type: Recreate stops all pods in a Brainstore role before creating replacements. This causes a complete role outage, and for the writer (which defaults to a single replica) pauses background processing until the replacement becomes ready.
Slow rollouts extend the period during which old and new Brainstore versions run together. Do not change brainstoreWalFooterVersion in the same upgrade as an image version bump, except where the documented data plane 2.0 upgrade sequence explicitly permits it.
The chart does not create PodDisruptionBudgets for Brainstore, so these rollout controls do not limit voluntary disruptions such as node drains or protect against involuntary pod or node failures.

Brainstore resource configuration

This section applies to GCP and Azure deployments using the Helm chart (v5.0.1+). AWS deployments manage Brainstore resources automatically.
Starting in Helm chart v5.0.1, the resources block for each Brainstore component (brainstore.reader, brainstore.writer, brainstore.fastreader) is passed through as-is to the Kubernetes pod spec. You can omit limits entirely, set them to {}, or supply any valid Kubernetes resource spec.
Omitting limits sets the pod QoS class to Burstable, which prevents CPU throttling and allows pods to use available node capacity. This can improve query performance on nodes with spare capacity, but increases the risk of resource contention if multiple pods compete for the same node.

Auto-derived Brainstore environment variables

As of Helm chart v5.1.0, BRAINSTORE_RESPONSE_CACHE_URI and BRAINSTORE_CODE_BUNDLE_URI are automatically populated from your objectStorage configuration and do not need to be set manually. The chart derives these values as follows: If you previously configured these via extraEnvVars, remove those overrides after upgrading to v5.1.0 to avoid conflicts.