我在尝试使用时遇到了问题%dopar%
and foreach()
与一个R6
班级。四处搜索,我只能找到两个与此相关的资源,一个未答复所以问题和一个开放的GitHub问题 on the R6
存储库。
在一条评论(即 GitHub 问题)中,建议通过重新分配parent_env
类的SomeClass$parent_env <- environment()
。我想了解到底是什么environment()
指的是这个表达式(即SomeClass$parent_env <- environment()
) 内被调用%dopar%
of foreach
?
这是一个最小的可重现示例:
Work <- R6::R6Class("Work",
public = list(
values = NULL,
initialize = function() {
self$values <- "some values"
}
)
)
现在,以下Task
类使用Work
构造函数中的类。
Task <- R6::R6Class("Task",
private = list(
..work = NULL
),
public = list(
initialize = function(time) {
private$..work <- Work$new()
Sys.sleep(time)
}
),
active = list(
work = function() {
return(private$..work)
}
)
)
In the Factory
类,该Task
类被创建并且foreach
实施于..m.thread()
.
Factory<- R6::R6Class("Factory",
private = list(
..warehouse = list(),
..amount = NULL,
..parallel = NULL,
..m.thread = function(object, ...) {
cluster <- parallel::makeCluster(parallel::detectCores() - 1)
doParallel::registerDoParallel(cluster)
private$..warehouse <- foreach::foreach(1:private$..amount, .export = c("Work")) %dopar% {
# What exactly does `environment()` encapsulate in this context?
object$parent_env <- environment()
object$new(...)
}
parallel::stopCluster(cluster)
},
..s.thread = function(object, ...) {
for (i in 1:private$..amount) {
private$..warehouse[[i]] <- object$new(...)
}
},
..run = function(object, ...) {
if(private$..parallel) {
private$..m.thread(object, ...)
} else {
private$..s.thread(object, ...)
}
}
),
public = list(
initialize = function(object, ..., amount = 10, parallel = FALSE) {
private$..amount = amount
private$..parallel = parallel
private$..run(object, ...)
}
),
active = list(
warehouse = function() {
return(private$..warehouse)
}
)
)
然后,它被称为:
library(foreach)
x = Factory$new(Task, time = 2, amount = 10, parallel = TRUE)
没有以下行object$parent_env <- environment()
,它会抛出一个错误(即,如其他两个链接中提到的):Error in { : task 1 failed - "object 'Work' not found"
.
我想知道,(1)分配时有哪些潜在的陷阱parent_env
inside foreach
(2)为什么它首先有效?
更新1:
- 我回来了
environment()
从内部foreach()
,使得private$..warehouse
捕捉这些环境
- using
rlang::env_print()
在调试会话中(即browser()
声明紧随其后foreach
已结束执行)它们的组成如下:
Browse[1]> env_print(private$..warehouse[[1]])
# <environment: 000000001A8332F0>
# parent: <environment: global>
# bindings:
# * Work: <S3: R6ClassGenerator>
# * ...: <...>
Browse[1]> env_print(environment())
# <environment: 000000001AC0F890>
# parent: <environment: 000000001AC20AF0>
# bindings:
# * private: <env>
# * cluster: <S3: SOCKcluster>
# * ...: <...>
Browse[1]> env_print(parent.env(environment()))
# <environment: 000000001AC20AF0>
# parent: <environment: global>
# bindings:
# * private: <env>
# * self: <S3: Factory>
Browse[1]> env_print(parent.env(parent.env(environment())))
# <environment: global>
# parent: <environment: package:rlang>
# bindings:
# * Work: <S3: R6ClassGenerator>
# * .Random.seed: <int>
# * Factory: <S3: R6ClassGenerator>
# * Task: <S3: R6ClassGenerator>