•1 min read•from Machine Learning
How is RLCD (jev) RL? [D]
Our take
The recent YouTube presentation on RLCD (jev) RL [D] has sparked considerable discussion, particularly regarding the role of Reinforcement Learning (RL). While jev's outputs—Choice, Score, and Noul—are readily differentiable using standard methods like cross-entropy or MSE, the integration of RL appears, to some, primarily for marketing purposes. A key question arises: what would a viable RL environment even look like within this framework?
Just saw the YouTube presentation and I was left wondering this question.
If jev only outputs Choice, Score, or Noul … well those are all perfectly differentiable. (Cross entropy or mse)
I don’t know if I’m missing something or if adding RL is just for marketing.
Like what would an RL environment even look like?
[link] [comments]
Read on the original site
Open the publisher's page for the full experience