# Access worker node environment variables in head node

**URL:** <https://discuss.ray.io/t/access-worker-node-environment-variables-in-head-node/9588>\
**Category:** Ray Core\
**Created:** [March 2, 2023, 5:20am UTC](https://discuss.ray.io/t/access-worker-node-environment-variables-in-head-node/9588 "2023-03-02T05:20:01Z")\
**Posts on this page:** 18\
**Page:** 1

<div class="post-metadata">

**Author:** ![shyampatel](https://avatars.discourse-cdn.com/v4/letter/s/cab0a1/32.png) [@shyampatel](https://discuss.ray.io/u/shyampatel)\
**Post date:** [March 2, 2023, 5:20am UTC](https://discuss.ray.io/t/access-worker-node-environment-variables-in-head-node/9588/1 "2023-03-02T05:20:01Z")

</div>

**How severe does this issue affect your experience of using Ray?**

- Medium: It contributes to significant difficulty to complete my task, but I can work around it.

I am using cluster.yaml file to create cluster. One of my requirement here is to access environment variable of worker node in head node.

---

<div class="post-metadata">

**Author:** ![jjyao](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/jjyao/32/1799_2.png) [@jjyao](https://discuss.ray.io/u/jjyao)\
**Post date:** [March 3, 2023, 5:16pm UTC](https://discuss.ray.io/t/access-worker-node-environment-variables-in-head-node/9588/2 "2023-03-03T17:16:39Z")

</div>

Could you give an example of what you want to do?

Can you set the same environment variables before starting head and worker nodes?

---

<div class="post-metadata">

**Author:** ![shyampatel](https://avatars.discourse-cdn.com/v4/letter/s/cab0a1/32.png) [@shyampatel](https://discuss.ray.io/u/shyampatel)\
**Post date:** [March 4, 2023, 12:00pm UTC](https://discuss.ray.io/t/access-worker-node-environment-variables-in-head-node/9588/3 "2023-03-04T12:00:42Z")

</div>

So, I am using cluster.yaml file for cluster creation and all the servers I am using as nodes are on-premise.

Now, some of the servers are with cuda GPU, some are with i5 processor, and some are with i7 processor. So I am defining an environment variable for **job capacity** based on the kind of CPU or GPU, a node has.

Now, to define the environment variables at the node level, I have written a shell script in each worker node **to start the worker node with defined variables** , and I am calling this script in the **worker node setup section in the cluster.yaml** file.

Now, one of my requirements is, I want to **read those environment variables in the head node** , which are defined in the worker node.

---

<div class="post-metadata">

**Author:** ![jjyao](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/jjyao/32/1799_2.png) [@jjyao](https://discuss.ray.io/u/jjyao)\
**Post date:** [March 6, 2023, 5:41pm UTC](https://discuss.ray.io/t/access-worker-node-environment-variables-in-head-node/9588/4 "2023-03-06T17:41:51Z")

</div>

So each worker is started with different env vars and you want to know the env vars for each worker from the head node? Like you want to get a map from worker node id to env vars? Is my understanding correct?

---

<div class="post-metadata">

**Author:** ![shyampatel](https://avatars.discourse-cdn.com/v4/letter/s/cab0a1/32.png) [@shyampatel](https://discuss.ray.io/u/shyampatel)\
**Post date:** [March 7, 2023, 5:53am UTC](https://discuss.ray.io/t/access-worker-node-environment-variables-in-head-node/9588/5 "2023-03-07T05:53:01Z")

</div>

> [@jjyao](#):
>
> Like you want to get a map from worker node id to env vars?

Exactly, I want the same.

---

<div class="post-metadata">

**Author:** ![jjyao](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/jjyao/32/1799_2.png) [@jjyao](https://discuss.ray.io/u/jjyao)\
**Post date:** [March 8, 2023, 10:43pm UTC](https://discuss.ray.io/t/access-worker-node-environment-variables-in-head-node/9588/6 "2023-03-08T22:43:17Z")

</div>

> [@shyampatel](#):
>
> So I am defining an environment variable for **job capacity** based on the kind of CPU or GPU, a node has.

Could you elaborate more on how these env vars will be used? I’m trying to see if you need env vars or Ray custom resources.

---

<div class="post-metadata">

**Author:** ![shyampatel](https://avatars.discourse-cdn.com/v4/letter/s/cab0a1/32.png) [@shyampatel](https://discuss.ray.io/u/shyampatel)\
**Post date:** [March 9, 2023, 3:54am UTC](https://discuss.ray.io/t/access-worker-node-environment-variables-in-head-node/9588/7 "2023-03-09T03:54:10Z")

</div>

@jjyao  
Actually, we have started with defining custom resources only, and it was working fine till the time. But now, for some new requirements, custom resources are restricting us to build generalised solution. With custom resources it’s possible, but it’s increasing complexity of our solution, which we don’t want. That’s why we come up with idea of using Environment variables.

> [@jjyao](#):
>
> Could you elaborate more on how these env vars will be used?

So, there are two main things which defines number of jobs a node can run: 1) memory (RAM) 2) type of CPU/GPU.

Now based on our async actor implementation, our single detection actor can handle multiple jobs. So, for any node, manually we will test that **how many instance** of detection actor it can handle in memory and **how many jobs per instance** of detection actor it can handle based on computation power. Based on efficient results from testing, finally we will have four environment variables, which can define #instance & #jobs\_per\_instance per CPU & GPU.

---

<div class="post-metadata">

**Author:** ![shyampatel](https://avatars.discourse-cdn.com/v4/letter/s/cab0a1/32.png) [@shyampatel](https://discuss.ray.io/u/shyampatel)\
**Post date:** [March 13, 2023, 4:12am UTC](https://discuss.ray.io/t/access-worker-node-environment-variables-in-head-node/9588/8 "2023-03-13T04:12:59Z")

</div>

@jjyao I hope I answered your question properly. Please let me know, if any further info I can provide.

---

<div class="post-metadata">

**Author:** ![jjyao](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/jjyao/32/1799_2.png) [@jjyao](https://discuss.ray.io/u/jjyao)\
**Post date:** [March 14, 2023, 4:18pm UTC](https://discuss.ray.io/t/access-worker-node-environment-variables-in-head-node/9588/9 "2023-03-14T16:18:11Z")

</div>

Sorry for the late reply. Let’s say now you know how many instances of detection actor you can run per node, how are you going to launch different number of actors to different node?

---

<div class="post-metadata">

**Author:** ![shyampatel](https://avatars.discourse-cdn.com/v4/letter/s/cab0a1/32.png) [@shyampatel](https://discuss.ray.io/u/shyampatel)\
**Post date:** [March 16, 2023, 11:08am UTC](https://discuss.ray.io/t/access-worker-node-environment-variables-in-head-node/9588/10 "2023-03-16T11:08:29Z")

</div>

I can find number of active instance of a actor based on **list\_actors()** state api, where I can group actors based on node ip. Now, based on capacity I can check, if I can run new instance or not. While creating actor, I am including info in actor name itself, Ex. **detactor\_ip\_address\_GPU** , which will help me for counting actor instances for CPU and GPU. Now to assign it to specific node, I am using custom resource **“node:ip\_address”**.

---

<div class="post-metadata">

**Author:** ![jjyao](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/jjyao/32/1799_2.png) [@jjyao](https://discuss.ray.io/u/jjyao)\
**Post date:** [March 16, 2023, 9:14pm UTC](https://discuss.ray.io/t/access-worker-node-environment-variables-in-head-node/9588/11 "2023-03-16T21:14:21Z")

</div>

I may not have the full picture but I still think we should use custom resources at least for **how many instances** of detection actor each node can run.

Also instead of using node ip to pin a task to node, you can use `NodeAffinityScheduingStrategy` ([Scheduling — Ray 2.3.0](https://docs.ray.io/en/releases-2.3.0/ray-core/scheduling/index.html#nodeaffinityschedulingstrategy)).

If you want to get worker’s env vars, I think you can launch a task to each node and returns the env vars.

---

<div class="post-metadata">

**Author:** ![Jules\_Damji](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/jules_damji/32/4058_2.png) [@Jules\_Damji](https://discuss.ray.io/u/Jules_Damji)\
**Post date:** [March 16, 2023, 10:04pm UTC](https://discuss.ray.io/t/access-worker-node-environment-variables-in-head-node/9588/12 "2023-03-16T22:04:29Z")

</div>

@shyampatel I hope @jjyao has provided you with sufficient guidance and answers. As pointed out, will the `NodeAffinitySchedulingStrategy` work for you?

---

<div class="post-metadata">

**Author:** ![shyampatel](https://avatars.discourse-cdn.com/v4/letter/s/cab0a1/32.png) [@shyampatel](https://discuss.ray.io/u/shyampatel)\
**Post date:** [March 17, 2023, 6:48am UTC](https://discuss.ray.io/t/access-worker-node-environment-variables-in-head-node/9588/13 "2023-03-17T06:48:38Z")

</div>

@Jules_Damji I have never used `NodeAffinitySchedulingStrategy`, so will have to go through the functionality and how I can integrate in our pipeline flow. But still, my requirement will be still needed to access environment variables of worker node.

As @jjyao mentioned, one method is I can create a task and assigned it to the respective node, which can return values of environment variables. Still, I am looking for an easy way to do this, if possible.

---

<div class="post-metadata">

**Author:** ![shyampatel](https://avatars.discourse-cdn.com/v4/letter/s/cab0a1/32.png) [@shyampatel](https://discuss.ray.io/u/shyampatel)\
**Post date:** [March 17, 2023, 6:52am UTC](https://discuss.ray.io/t/access-worker-node-environment-variables-in-head-node/9588/14 "2023-03-17T06:52:27Z")

</div>

@Jules_Damji one quick help required here. Can I assign specific actor to specific worker ip based on `NodeAffinitySchedulingStrategy` ?

---

<div class="post-metadata">

**Author:** ![jjyao](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/jjyao/32/1799_2.png) [@jjyao](https://discuss.ray.io/u/jjyao)\
**Post date:** [March 20, 2023, 7:42pm UTC](https://discuss.ray.io/t/access-worker-node-environment-variables-in-head-node/9588/15 "2023-03-20T19:42:16Z")

</div>

Yea, you can assign a specific actor to a specific worker `id` (not ip) based on `NodeAffinitySchedulingStrategy`

---

<div class="post-metadata">

**Author:** ![Jules\_Damji](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/jules_damji/32/4058_2.png) [@Jules\_Damji](https://discuss.ray.io/u/Jules_Damji)\
**Post date:** [March 20, 2023, 11:21pm UTC](https://discuss.ray.io/t/access-worker-node-environment-variables-in-head-node/9588/16 "2023-03-20T23:21:44Z")

</div>

@shyampatel Does that answer your question?

---

<div class="post-metadata">

**Author:** ![shyampatel](https://avatars.discourse-cdn.com/v4/letter/s/cab0a1/32.png) [@shyampatel](https://discuss.ray.io/u/shyampatel)\
**Post date:** [March 21, 2023, 4:16am UTC](https://discuss.ray.io/t/access-worker-node-environment-variables-in-head-node/9588/17 "2023-03-21T04:16:39Z")

</div>

Yes, That’s what I want to know. I have started going through the functionalities of `NodeAffinitySchedulingStrategy`, it will take time for me to update complete flow of my pipeline.

But, still with this also, it’s not solving my complete requirement which I have mentioned in below answer:

> [@shyampatel](#):
>
> So, there are two main things which defines number of jobs a node can run: 1) memory (RAM) 2) type of CPU/GPU.

Here, I have mentioned the use case of my environment variables:

> [@shyampatel](#):
>
> I can find number of active instance of a actor based on **list\_actors()** state api

---

<div class="post-metadata">

**Author:** ![jjyao](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/jjyao/32/1799_2.png) [@jjyao](https://discuss.ray.io/u/jjyao)\
**Post date:** [March 22, 2023, 6:26pm UTC](https://discuss.ray.io/t/access-worker-node-environment-variables-in-head-node/9588/18 "2023-03-22T18:26:59Z")

</div>

Ray doesn’t support getting environment variables of worker nodes so you need to do your own thing. One possibility is launching a task to each worker node to collect that.
