Styrow.dev
Question 1 of 2
Selenium Topic: Parallel Execution hard

How do you diagnose and resolve intermittent OutOfMemoryErrors and thread starvation in a large-scale Selenium-Java parallel execution suite?

Question

How do you diagnose and resolve intermittent OutOfMemoryErrors and thread starvation in a large-scale Selenium-Java parallel execution suite?

Answer

In a distributed CI/CD environment running a Selenium-Java suite with 50+ parallel test workers, engineers often encounter intermittent OutOfMemoryError: Java heap space or java.lang.OutOfMemoryError: unable to create new native thread failures. These failures are typically non-deterministic and do not reproduce in local development environments, which usually run with a smaller thread count or higher memory limits. This scenario presents a classic resource management dilemma where the interaction between the JVM thread pool, the Selenium WebDriver lifecycle, and the underlying OS process management creates a bottleneck.

The root cause is rarely the test logic itself but rather the accumulation of non-GC-able resources or thread leaks during the high-concurrency phase. In Java, WebDriver instances are not lightweight; they hold references to native browser processes (Chromedriver, Geckodriver) and internal state that requires explicit termination. If the test framework (e.g., JUnit 5, TestNG) or the custom executor does not guarantee the invocation of the cleanup logic (e.g., driver.quit()) upon every exception path, or if the garbage collector cannot keep up with the allocation rate of new driver contexts during peak load, the heap will fragment and exhaust.

Furthermore, the default ForkJoinPool or Executors implementations in Java may not be configured to handle the specific I/O-bound nature of browser automation. If the maximum pool size is too small relative to the number of parallel tests, threads will queue up, leading to timeouts. Conversely, if the pool size is too large without corresponding memory tuning, the JVM may attempt to spawn more threads than the OS allows, resulting in unable to create new native thread.

To resolve this, one must implement a multi-layered defense strategy:

  1. Explicit Thread Pool Management: Do not rely on default executors. Configure a ThreadPoolExecutor with a core and maximum pool size that matches the CI runner’s CPU cores and memory limits. For I/O-bound tasks like Selenium, the optimal thread count is often higher than the CPU count, but must be capped to prevent context-switching overhead and memory bloat.

  2. Rigid Resource Cleanup: Implement a @AfterEach or @AfterClass method that executes driver.quit() in a finally block. This ensures that even if a test throws an AssertionError or WebDriverException, the browser process is killed and the native memory is released. Additionally, clear the browser context (driver.manage().deleteAllCookies()) before quitting to prevent state leakage.

  3. JVM Memory Tuning: Adjust the JVM arguments in the CI build configuration. Increase the heap size (-Xmx) to accommodate the peak memory usage of parallel driver instances. Monitor the heap usage via JMX or logging to ensure that the GC is not thrashing.

  4. OS-Level Resource Limits: In containerized CI environments (e.g., Kubernetes, Docker), ensure that the container has sufficient CPU and memory limits. Selenium drivers are native processes and are not subject to the JVM’s memory management; they require OS-level resources. If the container limit is too low, the OS will kill the browser process, causing the driver to hang or throw a SessionNotCreatedException.

Here is a production-grade example of how to structure the test execution with proper resource management:

import org.junit.jupiter.api.AfterEach;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
import org.openqa.selenium.JavascriptExecutor;
import org.openqa.selenium.WebDriver;
import org.openqa.selenium.chrome.ChromeDriver;
import org.openqa.selenium.chrome.ChromeOptions;
import org.openqa.selenium.remote.RemoteWebDriver;
import org.openqa.selenium.remote.DesiredCapabilities;

import java.net.URL;
import java.time.Duration;

public class ParallelExecutionExample {

    private WebDriver driver;

    @BeforeEach
    void setUp() {
        ChromeOptions options = new ChromeOptions();
        options.addArguments("--headless", "--no-sandbox", "--disable-dev-shm-usage");
        
        // Ensure the driver is connected to a remote grid or local instance
        // In a real CI/CD environment, this URL would point to a Selenium Grid node
        try {
            driver = new RemoteWebDriver(new URL("http://localhost:4444/wd/hub"), options);
        } catch (Exception e) {
            throw new RuntimeException("Failed to create WebDriver", e);
        }
        
        // Set explicit timeouts to prevent indefinite hangs
        driver.manage().timeouts().implicitlyWait(Duration.ofSeconds(10));
        driver.manage().timeouts().pageLoadTimeout(Duration.ofSeconds(30));
    }

    @AfterEach
    void tearDown() {
        // CRITICAL: Always quit the driver in a finally-like manner
        // JUnit 5 @AfterEach is guaranteed to run even if the test fails
        if (driver != null) {
            try {
                driver.quit();
            } catch (Exception e) {
                // Log the error but do not fail the test due to cleanup issues
                System.err.println("Error quitting driver: " + e.getMessage());
            } finally {
                driver = null;
            }
        }
    }

    @Test
    void testParallelExecution() {
        driver.get("https://example.com");
        // Test logic here
        System.out.println("Title: " + driver.getTitle());
    }
}

For the parallel execution itself, using a custom ThreadPoolExecutor provides better control than ForkJoinPool:

import java.util.concurrent.*;

public class SeleniumTestExecutor {

    private static final int CORE_POOL_SIZE = 10;
    private static final int MAX_POOL_SIZE = 20;
    private static final long KEEP_ALIVE_TIME = 60L;

    private final ExecutorService executorService;

    public SeleniumTestExecutor() {
        this.executorService = new ThreadPoolExecutor(
                CORE_POOL_SIZE,
                MAX_POOL_SIZE,
                KEEP_ALIVE_TIME,
                TimeUnit.SECONDS,
                new LinkedBlockingQueue<>(100),
                new ThreadFactory() {
                    private int count = 0;
                    @Override
                    public Thread newThread(Runnable r) {
                        Thread t = new Thread(r, "Selenium-Worker-" + count++);
                        t.setDaemon(true);
                        return t;
                    }
                },
                new ThreadPoolExecutor.CallerRunsPolicy() // Backpressure handling
        );
    }

    public void executeTests(List<Runnable> tests) {
        List<Future<?>> futures = new ArrayList<>();
        for (Runnable test : tests) {
            futures.add(executorService.submit(test));
        }

        // Wait for all tests to complete
        for (Future<?> future : futures) {
            try {
                future.get();
            } catch (InterruptedException | ExecutionException e) {
                Thread.currentThread().interrupt();
                System.err.println("Test execution failed: " + e.getMessage());
            }
        }
    }

    public void shutdown() {
        executorService.shutdown();
        try {
            if (!executorService.awaitTermination(60, TimeUnit.SECONDS)) {
                executorService.shutdownNow();
            }
        } catch (InterruptedException e) {
            executorService.shutdownNow();
            Thread.currentThread().interrupt();
        }
    }
}

The key to solving this problem is not just writing the code correctly but monitoring the system. In a production CI/CD environment, you should integrate JVM metrics (heap usage, thread count) into your CI logs. If the heap usage spikes consistently before the OOM error, it indicates a memory leak in the test setup or a lack of proper cleanup. If the thread count spikes, it indicates a thread leak or an improperly configured pool. By combining explicit resource management, JVM tuning, and OS


📲 Practice Offline on Mobile: Download the free QA Automation & SDET Prep app on Google Play & App Store.