noreward-rl icon indicating copy to clipboard operation
noreward-rl copied to clipboard

extrinsic and intrinsic combination

Open murtazabasu opened this issue 5 years ago • 1 comments

Hello, I am trying to implement ICM in PPO with both extrinsic and intrinsic combination. I have seen in few repos where they weight out an extrinsic reward more than intrinsic i.e. combine_reward = (1-int_coef) * rewards + int_coef * intrinsic_reward whereint_coeff = 0.01which reduces the effect of intrinsic rewards significantly. Seeing your paper, you have nowhere mentioned this sort of equation for both the rewards. I wonder if you can tell me that the equation mentioned above can be implemented for a dual reward setting.

murtazabasu avatar Dec 23 '19 11:12 murtazabasu

Hello, do you understand the relationship between external rewards and internal rewards? how to adjust int_coef parameters.

Joll123 avatar May 24 '20 13:05 Joll123