Wiring Spring Boot Actuator health checks into a Grails application

Most production outages I have investigated in my years working with JVM teams in Sydney and Melbourne share the same root cause: nobody knew the service was broken until a customer in Brisbane complained. Health checks are the cheapest insurance policy a Grails shop can buy, and Spring Boot Actuator delivers almost all of that insurance for free because every Grails 5 application is already a Spring Boot 3 application under the hood. The framework exposes operational endpoints through the actuator module, and Groovy's terse syntax makes the small amount of custom code feel almost invisible.

Australian teams have a particular relationship with observability tooling because of the tyranny of distance. An application running in ap-southeast-2 with a sync replica in ap-southeast-4 can fail in ways that monitoring dashboards in the Northern Hemisphere miss during their business day. Wiring up proper health endpoints means a lone on-call developer in Perth at 2am AEST can look at a single JSON payload and decide whether to roll back or roll over and go back to sleep. There is no shame in wanting fewer pages.

The Spring Boot ecosystem treats health as a first-class concern, and Grails inherits the same plumbing through its dependency on spring-boot-starter-actuator. Once the dependency is on the classpath, the framework auto-configures a HealthEndpoint that aggregates information from any bean implementing the HealthIndicator interface. That aggregation model is the part that makes the integration feel native rather than bolted on.

This walkthrough shows how to turn on the actuator in a Grails 5 project, write a couple of Groovy health indicators that talk to a database and a downstream REST service, expose separate liveness and readiness probes for Kubernetes, and lock the endpoint down so it does not leak internal details to the public internet. Along the way I will point at a few resources that complement the topic, including a primer on mathematics historical roots for the curious reader, and a related Grails Example article on building-grails-custom-error-pages-for-different-http-status-codes that pairs well with custom health response shaping.

Why Spring Boot Actuator fits naturally with Grails

Grails 5 is built on Spring Boot 3, which means the actuator integration is not an add-on, it is a dependency you include in build.gradle. The autoconfiguration classes that ship with spring-boot-starter-actuator scan the application context at startup and register a /actuator/health endpoint, a /actuator/info endpoint, and a small set of management beans. Because Grails registers the same Spring application context that any Boot app would use, those beans are discovered without any custom configuration in the vast majority of cases.

The default health response is a simple {"status":"UP"} JSON object, which is fine for a smoke test but rarely what a team operating in production actually needs. Once you start adding data sources, message brokers, or external HTTP dependencies, you want the health endpoint to reflect the health of those collaborators, not just the health of the JVM. The HealthIndicator interface is the extension point that makes that possible, and Spring Boot already ships a dozen of them for common integrations like DataSource, DiskSpace, Redis, RabbitMQ, and Elasticsearch.

For Australian teams running on AWS Sydney, the disk space indicator alone has saved more than one deployment from running out of /var/lib space during a log retention blowout. The data source indicator saved a team I worked with in Adelaide from a much nastier incident: a primary-replica failover had happened overnight, and the health endpoint went from UP to DOWN with a clear message about the connection pool exhaustion. That information was in the logs, but the health endpoint made it impossible to ignore.

Enabling the actuator endpoints in a Grails 5 app

The shortest path to a working actuator is two lines in build.gradle: add the spring-boot-starter-actuator dependency and rebuild. By default, only the health endpoint is exposed over HTTP, which is exactly the conservative behaviour most teams want when they first turn the feature on. The other endpoints — metrics, env, beans, mappings, configprops — are all available on the management port but not exposed, so you can curl them from inside the pod without worrying about leaking them to a load balancer.

The classic second step is to flip the show-details property so the health endpoint reports per-component status instead of just an aggregate. In application.yml, that looks like:

management:
  endpoint:
    health:
      show-details: when_authorized
      show-components: always
  endpoints:
    web:
      exposure:
        include: health, info, metrics

The when_authorized value pairs with Spring Security to hide internal details from anonymous users while still returning the aggregate status. For a Grails app that already uses the spring-security-core plugin, a request from the load balancer probe gets {"status":"UP"} while an authenticated developer with the ROLE_ADMIN sees the full breakdown. That separation is exactly the pattern that Australian security teams at the big banks have been pushing for years, and it is the reason the default landed on when_authorized rather than always.

You can also change the management port so the actuator listens on a separate socket from the main application. That pattern is common in places like the Atlassian Sydney office where clusters share subnets and the ops team wants the actuator traffic isolated. management.server.port: 8081 and management.server.address: 127.0.0.1 will pin the actuator to localhost on a sidecar port that a sidecar proxy can scrape without exposing the endpoint to the wider VPC.

Crafting custom health indicators in Groovy

The real value of the actuator comes when you write your own HealthIndicator beans. The interface is a single method, Health health(), and Groovy's expressiveness makes the implementation feel like writing a sentence. A typical indicator for a downstream REST service might look like this:

@Component
class PaymentGatewayHealthIndicator implements HealthIndicator {
    private final RestClient restClient

    PaymentGatewayHealthIndicator(RestClient restClient) {
        this.restClient = restClient
    }

    @Override
    Health health() {
        try {
            def response = restClient.head('/health')
            if (response.status == 200) {
                return Health.up()
                    .withDetail('gateway', 'payments-prod')
                    .withDetail('latencyMs', response.elapsed)
                    .build()
            }
            return Health.down()
                .withDetail('statusCode', response.status)
                .build()
        } catch (Exception e) {
            return Health.down(e)
                .withDetail('gateway', 'payments-prod')
                .build()
        }
    }
}

A few details stand out. The Health.down(Exception) variant captures the exception class and message, which means a single glance at the JSON tells you whether the problem is a timeout, a connection refused, or an HTTP 503. The withDetail calls let you attach arbitrary key-value pairs that show up in the response body when show-details is enabled. For teams running across multiple regions — ap-southeast-2 for Sydney, ap-southeast-4 for Jakarta, sometimes ap-southeast-6 for Auckland — that detail block is where you record which region the indicator queried, because a down status in Jakarta should not page the on-call in Brisbane.

The same pattern works for any dependency: a DataSource, a S3 bucket, a Kafka topic, a Redis cluster. If your application can talk to it, you can write a health indicator for it, and the framework will aggregate them all into the response at /actuator/health. Aggregation behaviour is configurable through management.endpoint.health.group, which lets you bundle a subset of indicators into a named group that returns a single status — useful when you want to expose a liveness probe that only checks the JVM and a readiness probe that checks the JVM plus the database plus the downstream APIs.

Probes for liveness and readiness in container workloads

Kubernetes uses two probes, and they have very different jobs. The liveness probe tells the kubelet whether the container is alive — if it fails, Kubernetes restarts the pod. The readiness probe tells the kubelet whether the container is ready to receive traffic — if it fails, Kubernetes removes the pod from the service endpoints without restarting it. Conflating the two is a classic mistake that causes pod thrashing when a database hiccup takes down every replica at once.

Spring Boot 3 introduced the concepts of liveness state and readiness state as first-class beans, and the actuator wires them up automatically. In application.yml you map the probes to actuator endpoints:

management:
  endpoint:
    probes:
      enabled: true
  health:
    livenessstate:
      enabled: true
    readinessstate:
      enabled: true

The default liveness check returns UP as long as the application context is refreshed and the JVM is responsive, which is exactly the right behaviour — a stuck event loop or a deadlocked thread will cause the probe to time out and Kubernetes will restart the pod. The readiness check, on the other hand, runs your custom HealthIndicators and returns DOWN if any of them fail. While the database is unreachable, the pod is removed from the service but not restarted, and traffic is routed to the healthy replicas. Once the database comes back, the readiness check returns UP and the pod is added back to the rotation automatically.

For teams running on Amazon EKS in Sydney, the probe configuration lives in the deployment manifest and points at /actuator/health/liveness and /actuator/health/readiness respectively. The startup probe is also worth configuring for slow-starting Grails applications, because a fat JAR with a large Hibernate schema can take 30 to 60 seconds to come up and a default liveness probe will kill it before it finishes initialising. Setting initialDelaySeconds: 30, periodSeconds: 10, failureThreshold: 3 gives the application time to start without making the probe useless once it is running.

Securing and observing the health endpoint

The health endpoint is the most attacked surface on any internet-facing Spring Boot application, because every load balancer, scanner, and bot on the planet knows it exists. Returning too much information is a classic information leak: a health response that includes {"database":{"version":"14.10","pool":{"active":3}}} tells an attacker the exact database version and the current connection count, both useful for crafting a follow-on attack. The show-details: when_authorized setting is the right default, and pairing it with Spring Security rules that allow only ROLE_ACTUATOR to see details is the belt-and-braces approach.

For a Grails app using the spring-security-core plugin, the rule set looks like this in application.groovy:

grails.plugin.springsecurity.filterChain.chainMap = [
    '/actuator/**': 'ACTUATOR_FILTER',
    '/**': 'JOINED_FILTERS'
]

@Bean
ActuatorSecurityRules actuatorSecurityRules() {
    new ActuatorSecurityRules('/actuator/health/**', '/actuator/info')
}

The first entry lets the anonymous health check through — because Kubernetes needs to call it without credentials — while the second blocks the rest of the actuator behind authentication. If the cluster is using IAM roles for service accounts, the probe traffic comes from the kubelet and does not need a token. If the probe is hitting the endpoint from outside the cluster, terminate it at a load balancer that strips the path and forwards to a separate port.

Beyond security, the actuator's /actuator/metrics endpoint exposes the JVM, HTTP, and datasource metrics that Prometheus scrapers love. Pairing the actuator with Micrometer and a Prometheus registry turns a Grails application into a first-class citizen of the same observability stack the rest of the company uses, which matters for Australian teams who have already standardised on Grafana, Loki, and Tempo. A flat white and a Grafana dashboard are an increasingly common pairing in the cafes around Surry Hills and Richmond.

Operational patterns from Australian teams

Teams running Grails at scale in Australia tend to converge on a few habits. The first is to keep the default health response minimal at the load balancer and rely on the detailed endpoint behind a separate management port for the on-call engineer. The second is to write a health indicator for every external dependency the application cannot function without, and to make the indicator's withDetail output include the dependency name, the region, and the latency. The third is to wire the actuator metrics into the same alerting rules as the rest of the platform so a Grails service looks identical to a Node service in a dashboard.

A team at a fintech in Melbourne I worked with last summer had a particularly clean setup: every microservice exposed /actuator/health/readiness at a path that the ingress controller stripped and forwarded to the actuator port, while the liveness probe hit the same path on the application port. The detailed metrics were scraped by a Prometheus operator running in ap-southeast-2 and shipped to a central Thanos store. The whole stack was provisioned with Terraform, which meant a new region — ap-southeast-4 in Jakarta for a market expansion — took a single PR to add. The team's runbook for a DOWN status on the readiness probe was a single page that listed the custom health indicators, the SRE rotation, and the escalation path to Atlassian Statuspage if the outage lasted more than fifteen minutes.

The other habit worth copying is to use management.endpoint.health.group to define probe-specific subsets. A common pattern is a "liveness" group that only includes DiskSpaceHealthIndicator and a "readiness" group that includes DataSourceHealthIndicator, the custom downstream indicators, and the readiness state. That way the liveness probe only fails when the application is genuinely stuck, not when a single downstream API is having a bad day. The configuration is small but the effect on pod restart counts is significant, and any on-call engineer in Sydney at 3am will thank you for fewer unnecessary restarts.

If you have a Grails 5 application today, drop spring-boot-starter-actuator into build.gradle, flip show-details to when_authorized, and add a single custom HealthIndicator for the dependency that worries you the most. Rebuild, hit /actuator/health, and confirm the response includes the new component. From there, wire the liveness and readiness probes into your deployment manifest, secure the management port with Spring Security, and let the framework do the rest of the work. The next outage you would have missed will show up on the dashboard before a customer in Parramatta notices, and that is exactly the outcome you want. For more on shaping custom responses in Grails, the article on building-grails-custom-error-pages-for-different-http-status-codes covers the related pattern in detail. Subscribe to Grails Example for the next instalment, where I will walk through shipping these metrics into a Grafana Cloud workspace and writing the alerting rules that turn raw JSON into a page that actually wakes the right person up.