Oscillation of Episode Q0 during DDPG training

4 ビュー (過去 30 日間)

古いコメントを表示

Heesu Kim 2021 年 4 月 6 日

0
リンク

この質問への直接リンク

https://jp.mathworks.com/matlabcentral/answers/794607-oscillation-of-episode-q0-during-ddpg-training

コメント済み: Heesu Kim 2021 年 4 月 6 日

How do I interpret this kind of Episode Q0 oscillation?

The oscillation shows a pattern like up and down and the range also increases quite regularly.

According to other docs, they're saying the Q0 is supposed to approach actual discounted future reward as long as the critic network is designed properly.

Is this kind of Q0 oscillation just evidence that my critic network is not well-designed?

Is there any solution to work it out?

I'm not sure this question is acceptable to this community because I think it's more or less a theoretical issue.

1 件のコメント
-1 件の古いコメントを表示-1 件の古いコメントを非表示

Heesu Kim 2021 年 4 月 6 日

As a side note, I'm using DDPG + LSTM model that RL toolbox provides

サインインしてコメントする。

サインインしてこの質問に回答する。

回答 (0 件)

サインインしてこの質問に回答する。

カテゴリ

AI and Statistics Deep Learning Toolbox Applications Autonomous and Control Systems Reinforcement Learning

Help Center および File Exchange で Reinforcement Learning についてさらに検索

製品

Reinforcement Learning Toolbox

リリース

R2021a

Community Treasure Hunt

Find the treasures in MATLAB Central and discover how the community can help you!

Start Hunting!

Translated by

Oscillation of Episode Q0 during DDPG training

1 件のコメント
-1 件の古いコメントを表示-1 件の古いコメントを非表示

回答 (0 件)

参考

カテゴリ

タグ

製品

リリース

Community Treasure Hunt

Oscillation of Episode Q0 during DDPG training

1 件のコメント -1 件の古いコメントを表示-1 件の古いコメントを非表示

回答 (0 件)

参考

カテゴリ

タグ

製品

リリース

Community Treasure Hunt

1 件のコメント
-1 件の古いコメントを表示-1 件の古いコメントを非表示