1 min readfrom Machine Learning

How is RLCD (jev) RL? [D]

Our take

The recent YouTube presentation on RLCD (jev) RL [D] has sparked considerable discussion, particularly regarding the role of Reinforcement Learning (RL). While jev's outputs—Choice, Score, and Noul—are readily differentiable using standard methods like cross-entropy or MSE, the integration of RL appears, to some, primarily for marketing purposes. A key question arises: what would a viable RL environment even look like within this framework?

Just saw the YouTube presentation and I was left wondering this question.
If jev only outputs Choice, Score, or Noul … well those are all perfectly differentiable. (Cross entropy or mse)
I don’t know if I’m missing something or if adding RL is just for marketing.
Like what would an RL environment even look like?

submitted by /u/Relative_Wallaby_823
[link] [comments]

Read on the original site

Open the publisher's page for the full experience

View original article