# Worker nodes are IDLE

**URL:** <https://discuss.ray.io/t/worker-nodes-are-idle/4783>\
**Category:** Uncategorized\
**Created:** [January 25, 2022, 10:03am UTC](https://discuss.ray.io/t/worker-nodes-are-idle/4783 "2022-01-25T10:03:09Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![mataney](https://avatars.discourse-cdn.com/v4/letter/m/ac8455/32.png) [@mataney](https://discuss.ray.io/u/mataney)\
**Post date:** [January 25, 2022, 10:03am UTC](https://discuss.ray.io/t/worker-nodes-are-idle/4783/1 "2022-01-25T10:03:09Z")

</div>

Hi there, thanks for the great job!

I have a task that runs a quick inference of an ML model over the gpu. for this I set `min_workers: 20` for my worker node. After ray initalizes all the worker nodes almost all of the stay idle (usually only 2-4 out of the 20 is not idle).

What seem to be the problem?  
I also noticed that the head node is doing most of the workload, how can I do proper load balancing over all my nodes?  
Thanks.

Thanks my config:

```auto
cluster_name: gpucluster
max_workers: 100
upscaling_speed: 2.0
idle_timeout_minutes: 10
docker:
   image: "rayproject/ray:latest-gpu"
   container_name: "ray_container"

provider:
    type: gcp
    region: ...
    availability_zone: ...
    project_id: ...
auth:
    ssh_user: ray
available_node_types:
    head_node:
        min_workers: 0
        max_workers: 0
        resources: {"CPU": 4, "GPU": 1}
        node_config:
            machineType: n1-highmem-4
            tags:
              - items: ["allow-all"]
            disks:
              - boot: true
                autoDelete: true
                type: PERSISTENT
                initializeParams:
                  diskSizeGb: 100
                  sourceImage: projects/deeplearning-platform-release/global/images/family/common-cu113
            guestAccelerators:
              - acceleratorType: .../nvidia-tesla-p100
                acceleratorCount: 1
            metadata:
              items:
                - key: install-nvidia-driver
                  value: "True"
            scheduling:
              - onHostMaintenance: "terminate"
              - automaticRestart: true
    worker_node:
        min_workers: 20
        resources: {"CPU": 4, "GPU": 1}
        node_config:
            machineType: n1-highmem-4
            tags:
              - items: ["allow-all"]
            disks:
              - boot: true
                autoDelete: true
                type: PERSISTENT
                initializeParams:
                  diskSizeGb: 100
                  sourceImage: projects/deeplearning-platform-release/global/images/family/common-cu113
            scheduling:
              - preemptible: false
            guestAccelerators:
              - acceleratorType: .../nvidia-tesla-p100
                acceleratorCount: 1
            metadata:
              items:
                - key: install-nvidia-driver
                  value: "True"
            scheduling:
              - onHostMaintenance: "terminate"
              - automaticRestart: true

head_node_type: head_node

```

---

<div class="post-metadata">

**Author:** ![matthewdeng](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/matthewdeng/32/1446_2.png) [@matthewdeng](https://discuss.ray.io/u/matthewdeng)\
**Post date:** [January 26, 2022, 5:26am UTC](https://discuss.ray.io/t/worker-nodes-are-idle/4783/2 "2022-01-26T05:26:51Z")

</div>

Hey @mataney,

Would you be able to share what your Python script looks like? How are you spawning your models? Logically if you are running actors or tasks that require 1 GPU each in parallel, they should be run on different nodes (as each node has 1 GPU resource).
