Continual RL Workshop - Resources
1 directory member surfaced this signal.
Join free with your email
Free. No spam, ever — we'll never share your email address and you can opt out at any time. Already a subscriber? Log in
What credible people across AI noticed, why it matters, and where the field is converging or disagreeing.
One card per development. Sources are clustered; expert reactions remain attributable.
1 directory member surfaced this signal.
1 directory member surfaced this signal.
“There are a few more recent work, but as far as I recall, they are still under restrictive assumptions. There is a new paper came last month, but I haven't read it yet. Convergence of Monte Carlo Optimistic Policy Iteration: Beyond Uniform State-Action Upda…”
1 directory member surfaced this signal.
“There is John Tsitsiklis's work from 2002, but that's for the initial state update model, which feels sample inefficient. Moreover, it assumes either synchronous updates (all states are updated at the same time) or asynchronous but random with uniform distr…”