Why are policy gradients popular in RL when there exists a dual LP formulation in terms of occupation measures that can be solved easily?
AI Stack Exchange··Updated just now·29 sightings
AI brief
A question on AI Stack Exchange asks why policy gradient methods are popular in reinforcement learning when a dual linear programming formulation in terms of occupation measures can be solved easily.
Written by AI from AI Stack Exchange's published text. Read the original for full details.