It's really a method with only one enter, scenario, and only one output, action (or behavior) a. You can find neither a separate reinforcement input nor an guidance enter through the natural environment. The backpropagated value (secondary reinforcement) will be the emotion towards the consequence p… Read More