# Consistency of RNG\_BEGIN and RNG\_END

**URL:** <https://igraph.discourse.group/t/consistency-of-rng-begin-and-rng-end/641>\
**Category:** Development\
**Tags:** C, R\
**Created:** [24 February 2021 14:32 UTC](https://igraph.discourse.group/t/consistency-of-rng-begin-and-rng-end/641 "2021-02-24T14:32:40Z")\
**Posts on this page:** 18\
**Page:** 1

<div class="post-metadata">

**Author:** ![GroteGnoom](https://yyz2.discourse-cdn.com/free1/user_avatar/igraph.discourse.group/grotegnoom/32/428_2.png) [@GroteGnoom](https://igraph.discourse.group/u/GroteGnoom)\
**Post date:** [24 February 2021 14:32 UTC](https://igraph.discourse.group/t/consistency-of-rng-begin-and-rng-end/641/1 "2021-02-24T14:32:40Z")

</div>

Should `RNG_BEGIN` and `RNG_END` be put around every function call that call other functions that use randomness?

So if A calls B and B calls C and C calls RNG\_INTEGER, should A have

```auto
RNG_BEGIN();
B();
RNG_END();

```

and B have

```auto
RNG_BEGIN();
C();
RNG_END();

```

and then C

```auto
RNG_BEGIN();
RND_INTEGER(0, n);
RNG_END();

```

?

That seems what happens in practice sometimes, but at least for the non-R definitions of `RNG_BEGIN` and `RNG_END` this doesn’t seem to be necessary.

---

<div class="post-metadata">

**Author:** ![vtraag](https://yyz2.discourse-cdn.com/free1/user_avatar/igraph.discourse.group/vtraag/32/38_2.png) [@vtraag](https://igraph.discourse.group/u/vtraag)\
**Post date:** [24 February 2021 14:56 UTC](https://igraph.discourse.group/t/consistency-of-rng-begin-and-rng-end/641/2 "2021-02-24T14:56:14Z")

</div>

Good question! It actually begs the question if the nesting is doing things properly.

That is, if we effectively have two calls to `RNG_BEGIN()` after each other, it seems to imply that the seed is being read multiple time (which seems to be what `GetRNGState` does). This might effectively mean that the RNG is always reset to the same seed whenever `RNG_BEGIN()` is being called. The resulting state is only being written back by `PutRNGState`, which is done by `RNG_END()`. If this interpretation is correct, having these nested calls then effectively reduce the “randomness” substantially. I am not sure if this interpretation is correct though, perhaps @Gabor can weigh in indeed.

Perhaps it would be better if the `GetRNGstate()` and `PutRNGstate()` would actually only be present in the R-C interface glue code. This is also what seems to be done by `Rcpp`, if I understand correctly.

---

<div class="post-metadata">

**Author:** ![szhorvat](https://yyz2.discourse-cdn.com/free1/user_avatar/igraph.discourse.group/szhorvat/32/3_2.png) [@szhorvat](https://igraph.discourse.group/u/szhorvat)\
**Post date:** [24 February 2021 14:58 UTC](https://igraph.discourse.group/t/consistency-of-rng-begin-and-rng-end/641/3 "2021-02-24T14:58:57Z")

</div>

Here are the docs for these functions, but I don’t fully understand what they do and why they are needed:

[https://cran.r-project.org/doc/manuals/r-release/R-exts.html#Random-numbers](https://cran.r-project.org/doc/manuals/r-release/R-exts.html#Random-numbers)

---

<div class="post-metadata">

**Author:** ![szhorvat](https://yyz2.discourse-cdn.com/free1/user_avatar/igraph.discourse.group/szhorvat/32/3_2.png) [@szhorvat](https://igraph.discourse.group/u/szhorvat)\
**Post date:** [24 February 2021 15:00 UTC](https://igraph.discourse.group/t/consistency-of-rng-begin-and-rng-end/641/4 "2021-02-24T15:00:31Z")

</div>

> [@vtraag](#):
>
> This might effectively mean that the RNG is always reset to the same seed whenever `RNG_BEGIN()` is being called.

I don’t think that’s what it does (if it did, things would be broken), see here:

[https://stat.ethz.ch/pipermail/r-help/2004-March/046967.html](https://stat.ethz.ch/pipermail/r-help/2004-March/046967.html)

But I don’t understand the purpose of these R functions.

---

<div class="post-metadata">

**Author:** ![vtraag](https://yyz2.discourse-cdn.com/free1/user_avatar/igraph.discourse.group/vtraag/32/38_2.png) [@vtraag](https://igraph.discourse.group/u/vtraag)\
**Post date:** [24 February 2021 15:06 UTC](https://igraph.discourse.group/t/consistency-of-rng-begin-and-rng-end/641/5 "2021-02-24T15:06:07Z")

</div>

> [@szhorvat](#):
>
> I don’t think that’s what it does (if it did, things would be broken), see here:
> 
> [[R] generating normal numbers: GetRNGstate, PutRNGstate](https://stat.ethz.ch/pipermail/r-help/2004-March/046967.html)
> 
> But I don’t understand the purpose of these R functions.

I don’t believe this mail addresses my question. The approaches therein are indeed equivalent, if you interpret each `GetRNGstate` and `PutRNGstate` to effectuate a reading / setting of a seed. However, if you do this:

```C
GetRNGstate();
for(i=1; i < 1000; i++){
  x[i]=norm_rand();
}
// This below possibly resets the seed
GetRNGstate();
for(i=1; i < 1000; i++){
  y[i]=norm_rand();
}
PutRNGstate();

```

You would obtain twice the same random numbers, and `x == y`, if my interpretation is correct.

The purpose of `GetRNGstate` and `PutRNGstate` seems to be to ensure that every call to the R random number generator is consistent. Why that state is not simply being kept internally I don’t really understand either.

---

<div class="post-metadata">

**Author:** ![szhorvat](https://yyz2.discourse-cdn.com/free1/user_avatar/igraph.discourse.group/szhorvat/32/3_2.png) [@szhorvat](https://igraph.discourse.group/u/szhorvat)\
**Post date:** [24 February 2021 15:10 UTC](https://igraph.discourse.group/t/consistency-of-rng-begin-and-rng-end/641/6 "2021-02-24T15:10:23Z")

</div>

> [@vtraag](#):
>
> The purpose of `GetRNGstate` and `PutRNGstate` seems to be to ensure that every call to the R random number generator is consistent. Why that state is not simply being kept internally I don’t really understand either.

I might be wrong, but I think the purpose of `GetRNGState()` and `PutRNGState()` is the following:

The RNG is implemented in C, and thus its state is kept in a memory location that is easily accessible from C. But the RNG state is supposed to represented by the `.Random.seed` (misleading name—it’s not a seed but the state) variable in R. These functions simply ensure that the “C variable” and the “R variable” representing the state are always in sync.

Do you agree @vtraag?

---

<div class="post-metadata">

**Author:** ![vtraag](https://yyz2.discourse-cdn.com/free1/user_avatar/igraph.discourse.group/vtraag/32/38_2.png) [@vtraag](https://igraph.discourse.group/u/vtraag)\
**Post date:** [24 February 2021 15:13 UTC](https://igraph.discourse.group/t/consistency-of-rng-begin-and-rng-end/641/7 "2021-02-24T15:13:04Z")

</div>

> [@szhorvat](#):
>
> Do you agree @vtraag?

Yes, this is what I meant. So, if the `.Random.seed` is always read upon `GetRNGstate`, and sets the current state of the RNG in C, it effectively always resets the state.

---

<div class="post-metadata">

**Author:** ![szhorvat](https://yyz2.discourse-cdn.com/free1/user_avatar/igraph.discourse.group/szhorvat/32/3_2.png) [@szhorvat](https://igraph.discourse.group/u/szhorvat)\
**Post date:** [24 February 2021 15:15 UTC](https://igraph.discourse.group/t/consistency-of-rng-begin-and-rng-end/641/8 "2021-02-24T15:15:17Z")

</div>

The consequences are that for as long as it’s guaranteed that `RNG_BEGIN()` and `RNG_END()` always come in a pair (i.e. `RNG_END()` is guaranteed to execute whenever `RNG_BEGIN()` does, unless there is an error) we are fine.

Then the guideline for using these would be:

Use `RNG_BEGIN` and `RNG_END` only in functions that directly access the random number generator, i.e. call `igraph_rng_` functions (possibly through macros).

> [@GroteGnoom](#):
>
> So if A calls B and B calls C and C calls RNG\_INTEGER, should A have
> 
> ```auto
> RNG_BEGIN();
> B();
> RNG_END();
> 
> ```
> 
> and B have
> 
> ```auto
> RNG_BEGIN();
> C();
> RNG_END();
> 
> ```
> 
> and then C
> 
> ```auto
> RNG_BEGIN();
> RND_INTEGER(0, n);
> RNG_END();
> 
> ```
> 
> ?

In this example the answer would be to only use `RNG_BEGIN` and `RNG_END` in function `C` (although using them in `A` and `B` won’t cause problems, except for some wasted CPU cycles).

---

<div class="post-metadata">

**Author:** ![vtraag](https://yyz2.discourse-cdn.com/free1/user_avatar/igraph.discourse.group/vtraag/32/38_2.png) [@vtraag](https://igraph.discourse.group/u/vtraag)\
**Post date:** [24 February 2021 15:22 UTC](https://igraph.discourse.group/t/consistency-of-rng-begin-and-rng-end/641/9 "2021-02-24T15:22:13Z")

</div>

> [@szhorvat](#):
>
> The consequences are that for as long as it’s guaranteed that `RNG_BEGIN()` and `RNG_END()` always come in a pair (i.e. `RNG_END()` is guaranteed to execute whenever `RNG_BEGIN()` does, unless there is an error) we are fine.

Yes, I agree. However, I am not sure if this is always guaranteed throughout the code base. For example, if you would do something like this:

```auto
RNG_BEGIN();

x = RNG_INTEGER(0, 100);

igraph_vector_shuffle(&v);

RNG_END();

```

the random state will get reset in `igraph_vector_shuffle` because it calls `RNG_BEGIN`. Hence, this “reduces the randomness” in that sense.

---

<div class="post-metadata">

**Author:** ![vtraag](https://yyz2.discourse-cdn.com/free1/user_avatar/igraph.discourse.group/vtraag/32/38_2.png) [@vtraag](https://igraph.discourse.group/u/vtraag)\
**Post date:** [24 February 2021 15:23 UTC](https://igraph.discourse.group/t/consistency-of-rng-begin-and-rng-end/641/10 "2021-02-24T15:23:43Z")

</div>

All in all, I think the safest option is to always _only_ get the random state for R on the entry point in C, which is done in the R-C glue code in the R interface, and again save the random state upon return in R.

---

<div class="post-metadata">

**Author:** ![szhorvat](https://yyz2.discourse-cdn.com/free1/user_avatar/igraph.discourse.group/szhorvat/32/3_2.png) [@szhorvat](https://igraph.discourse.group/u/szhorvat)\
**Post date:** [24 February 2021 15:26 UTC](https://igraph.discourse.group/t/consistency-of-rng-begin-and-rng-end/641/11 "2021-02-24T15:26:16Z")

</div>

So then all of this RNG\_BEGIN and RNG\_END is mostly useless leftover then?

---

<div class="post-metadata">

**Author:** ![vtraag](https://yyz2.discourse-cdn.com/free1/user_avatar/igraph.discourse.group/vtraag/32/38_2.png) [@vtraag](https://igraph.discourse.group/u/vtraag)\
**Post date:** [24 February 2021 15:29 UTC](https://igraph.discourse.group/t/consistency-of-rng-begin-and-rng-end/641/12 "2021-02-24T15:29:40Z")

</div>

> [@szhorvat](#):
>
> So then all of this RNG\_BEGIN and RNG\_END is mostly useless leftover then?

Well, yes, if we move all those calls to the R interface. However, before drawing that conclusion, perhaps we should wait for @Gabor to weigh in. Perhaps there are some other difficulties that we may have not considered yet.

Of course we do need to initialize the default RNG somehow, but that is a separate question, and might be done in a different way.

---

<div class="post-metadata">

**Author:** ![vtraag](https://yyz2.discourse-cdn.com/free1/user_avatar/igraph.discourse.group/vtraag/32/38_2.png) [@vtraag](https://igraph.discourse.group/u/vtraag)\
**Post date:** [24 February 2021 15:41 UTC](https://igraph.discourse.group/t/consistency-of-rng-begin-and-rng-end/641/13 "2021-02-24T15:41:58Z")

</div>

I have checked this with some small C code:

```auto
#include <R.h>

void rng_state(double* x1, double* x2)
{
    GetRNGstate();
    *x1 = unif_rand();
    GetRNGstate();
    *x2 = unif_rand();
    PutRNGstate();
}

```

You can compile this using `R CMD SHLIB test_rng.c`. Then fire up `R` and calling this function yields:

```auto
> dyn.load("test_rng.so")
> .C("rng_state", x1=0.0, x2=0.0)
$x1
[1] 0.1239771

$x2
[1] 0.7680124

> .C("rng_state", x1=0.0, x2=0.0)
$x1
[1] 0.46593

$x2
[1] 0.46593

> .C("rng_state", x1=0.0, x2=0.0)
$x1
[1] 0.9483618

$x2
[1] 0.9483618

> .C("rng_state", x1=0.0, x2=0.0)
$x1
[1] 0.6552871

$x2
[1] 0.6552871

```

As you can see, on the first call, the state is set for the first time, so the two results are not identical. However, any subsequent calls actually do lead to identical results. This shows that there might be potentials problems with the `RNG_BEGIN` and `RNG_END` calls for the R interface that “reduce the randomness”.

---

<div class="post-metadata">

**Author:** ![tamas](https://yyz2.discourse-cdn.com/free1/user_avatar/igraph.discourse.group/tamas/32/1261_2.png) [@tamas](https://igraph.discourse.group/u/tamas)\
**Post date:** [24 February 2021 18:16 UTC](https://igraph.discourse.group/t/consistency-of-rng-begin-and-rng-end/641/14 "2021-02-24T18:16:21Z")

</div>

In general, I think we shouldn’t try to match the semantics of igraph’s `RNG_BEGIN()` and `RNG_END()` to what R expects because there are other higher-level interfaces to igraph, which might need different semantics. Or they might not need it at all. @szhorvat, do you use these functions in the Mathematica interface? The Python interface does not make any use of them.

Anyway, I think the expected behaviour of these functions (I mean, what an ordinary user of the library would expect) should allow nested calls. There should be a counter that starts from zero and is increased with every call to `RNG_BEGIN()` and decreased with every call to `RNG_END()`. R’s `GetRNGState()` should be called when the counter goes from 0 to 1, and `PutRNGState()` should be called when the counter goes from 1 to 0. igraph should `abort()` when `RNG_END()` is called while the counter is zero as this indicates incorrect nesting. It is an open question whether the counter should be thread-local (maybe R’s RNG is not thread-safe?) or not.

---

<div class="post-metadata">

**Author:** ![szhorvat](https://yyz2.discourse-cdn.com/free1/user_avatar/igraph.discourse.group/szhorvat/32/3_2.png) [@szhorvat](https://igraph.discourse.group/u/szhorvat)\
**Post date:** [24 February 2021 18:34 UTC](https://igraph.discourse.group/t/consistency-of-rng-begin-and-rng-end/641/15 "2021-02-24T18:34:13Z")

</div>

> [@tamas](#):
>
> @szhorvat, do you use these functions in the Mathematica interface?

No, I don’t use them.

---

<div class="post-metadata">

**Author:** ![vtraag](https://yyz2.discourse-cdn.com/free1/user_avatar/igraph.discourse.group/vtraag/32/38_2.png) [@vtraag](https://igraph.discourse.group/u/vtraag)\
**Post date:** [24 February 2021 18:59 UTC](https://igraph.discourse.group/t/consistency-of-rng-begin-and-rng-end/641/16 "2021-02-24T18:59:33Z")

</div>

I think the general question is why we have them at all. They mostly seem to serve the R interface, and arguably, dealing with this should be done inside the R interface if possible, not in the C library. The only other usecase is to initialize the random number generator, but there might be other ways of dealing with this. Frankly, if we could simply get rid of the use of `RNG_BEGIN` and `RNG_END` all together, this would be my preference.

---

<div class="post-metadata">

**Author:** ![tamas](https://yyz2.discourse-cdn.com/free1/user_avatar/igraph.discourse.group/tamas/32/1261_2.png) [@tamas](https://igraph.discourse.group/u/tamas)\
**Post date:** [24 February 2021 19:41 UTC](https://igraph.discourse.group/t/consistency-of-rng-begin-and-rng-end/641/17 "2021-02-24T19:41:03Z")

</div>

One use-case that I could think of is lazy auto-seeding of the RNG from the current time, which can be done in `RNG_BEGIN()`. Actually, this also seems to be a side effect of `GetRNGstate()` in R, judging from the [source code](https://github.com/wch/r-source/blob/trunk/src/main/RNG.c#L404) – note that `GetRNGstate()` may call `Randomize()`, which sets the random seed. This can also explain why Vincent saw that the first call of his test code produced different values for `x1` and `x2` but not subsequent ones – the first call to `GetRNGstate()` that copied the RNG state from R-world to C-world copied an unseeded RNG so `GetRNGstate()` seeded the RNG, then the next call to `GetRNGstate()` copied the RNG state _again_, making the RNG unseeded, and then it re-seeded the RNG again. From the second call onwards, the base state was seeded so the results were identical.

If you want to lift `RNG_BEGIN()` and `RNG_END()` to the R-to-C glue code, then the glue code must either know whether the igraph function being called will use the RNG and prepare the RNG state accordingly, or must assume that _all_ igraph functions can potentially use the RNG from R, and call `GetRNGstate()` or `PutRNGstate()` unconditionally. Even when doing so, we cannot prevent nested calls to `RNG_BEGIN()` and `RNG_END()`: think about an igraph function that takes a callback function written in R – the callback function may call back to igraph via the glue code again, leading to a potential nesting.

So, my opinion so far, based on my limited understanding is that:

- it seems hard to delegate the responsibility of knowing when the RNG will be called (and thus the responsibiity of calling `GetRNGstate()` and `PutRNGstate()` to the R-to-C glue code.

- even if we did so, we cannot prevent nested calls so the counter-based approach that I outlined is probably still necessary

---

<div class="post-metadata">

**Author:** ![vtraag](https://yyz2.discourse-cdn.com/free1/user_avatar/igraph.discourse.group/vtraag/32/38_2.png) [@vtraag](https://igraph.discourse.group/u/vtraag)\
**Post date:** [25 February 2021 08:05 UTC](https://igraph.discourse.group/t/consistency-of-rng-begin-and-rng-end/641/18 "2021-02-25T08:05:18Z")

</div>

> [@tamas](#):
>
> One use-case that I could think of is lazy auto-seeding of the RNG from the current time, which can be done in `RNG_BEGIN()` .

Yes, this is exactly what is the only other use case that is currently present. However, the lazy auto-seeding does not need to be in there. To have these `RNG_BEGIN` (and unnecessary `RNG_END`) calls, just so that a user does not have to seed the RNG seems a bit overkill to me. It just requires a user to specify a single call to seed it. For higher level interfaces this can always be done (or does not need to be done), so most users will not even know the difference.

> [@tamas](#):
>
> If you want to lift `RNG_BEGIN()` and `RNG_END()` to the R-to-C glue code, then the glue code must either know whether the igraph function being called will use the RNG and prepare the RNG state accordingly, or must assume that _all_ igraph functions can potentially use the RNG from R, and call `GetRNGstate()` or `PutRNGstate()` unconditionally. Even when doing so, we cannot prevent nested calls to `RNG_BEGIN()` and `RNG_END()` : think about an igraph function that takes a callback function written in R – the callback function may call back to igraph via the glue code again, leading to a potential nesting.

Yes, all functions will simply have an `RNG_BEGIN` and `RNG_END` in this scenario. Callback functions need to be looked at carefully indeed, because they should _not_ have this in there. However, most callback functions will not be entry points into the C from R, so I doubt that this will become a problem. This needs to be double checked though.

> [@tamas](#):
>
> even if we did so, we cannot prevent nested calls so the counter-based approach that I outlined is probably still necessary

The solution that you propose does solve the problem indeed, regardless of whether the nesting is correct or not. Nevertheless, nested calls in themselves are in principle not a problem, it is the fact that there may be something happening prior to the nested call.

Let me check the R interface to see if there are such potential problems with callbacks that cannot be solved.
