Models & research
OpenAI
Variance reduction for policy gradient with action-dependent factorized baselines
This source did not provide an excerpt. Read the announcement at the publisher.
Excerpt supplied by the publisher’s feed