critic

Design and implementation of an adaptive critic-based neuro-fuzzy controller on an unmanned bicycle

Journal: :CoRR 2017

Ali Shafiekhani Mohammad J. Mahjoob Mehdi Akraminia

Abstract: Fuzzy critic-based learning forms a reinforcement learning method based on dynamic programming. In this paper, an adaptive critic-based neuro-fuzzy system is presented for an unmanned bicycle. The only information available for the critic agent is the system feedback which is interpreted as the last action performed by the controller in the previous state. The signal produced by the c...

متن کامل

Findings of the Panel of Psychological Inquiry Convened at Saint Michael’s College, May 13, 2008: The Case of “Anna”

2011

RONALD B. MILLER MARC KESSLER MARION BAUER SANDRA HOWELL KENNETH KREILING

This paper briefly describes the proceedings of the Panel of Inquiry held May 13, 2008 at Saint Michael’s College on the case of “Anna" (Podetz, 2008, 2011). It summarizes the advocate's and critic's positions on four claims and one counter-claim. The five judges independently voted to accept all four of the advocate’s claims (by votes of 5-0 or 4-1), and rejected the critic's counterclaim by a...

متن کامل

Online Learning of Optimal Control Solutions Using Integral Reinforcement Learning and Neural Networks

2011

Kyriakos G. Vamvoudakis Draguna Vrabie Frank L. Lewis

In this paper we introduce an online algorithm that uses integral reinforcement knowledge for learning the continuous-time optimal control solution for nonlinear systems with infinite horizon costs and partial knowledge of the system dynamics. This algorithm is a data based approach to the solution of the Hamilton-Jacobi-Bellman equation and it does not require explicit knowledge on the system’...

متن کامل

Model-free Adaptive Dynamic Programming for Optimal Control of Discrete-time Affine Nonlinear System ⋆

2014

Zhongpu Xia Dongbin Zhao Huajin Tang

In this paper, a model-free and effective approach is proposed to solve infinite horizon optimal control problem for affine nonlinear systems based on adaptive dynamic programming technique. The developed approach, referred to as the actor-critic structure, employs two multilayer perceptron neural networks to approximate the state-action value function and the control policy, respectively. It u...

متن کامل

An RLS-Based Natural Actor-Critic Algorithm for Locomotion of a Two-Linked Robot Arm

2005

Jooyoung Park Jongho Kim Daesung Kang

Recently, actor-critic methods have drawn much interests in the area of reinforcement learning, and several algorithms have been studied along the line of the actor-critic strategy. This paper studies an actor-critic type algorithm utilizing the RLS(recursive least-squares) method, which is one of the most efficient techniques for adaptive signal processing, together with natural policy gradien...

متن کامل

Sustainable ℓ2-regularized actor-critic based on recursive least-squares temporal difference learning

2017

Luntong Li Dazi Li Tianheng Song

Least-squares temporal difference learning (LSTD) has been used mainly for improving the data efficiency of the critic in actor-critic (AC). However, convergence analysis of the resulted algorithms is difficult when policy is changing. In this paper, a new AC method is proposed based on LSTD under discount criterion. The method comprises two components as the contribution: (1) LSTD works in an ...

متن کامل

Actor-critic models of the basal ganglia: new anatomical and computational perspectives

Journal: :Neural networks : the official journal of the International Neural Network Society 2002

Daphna Joel Yael Niv Eytan Ruppin

A large number of computational models of information processing in the basal ganglia have been developed in recent years. Prominent in these are actor-critic models of basal ganglia functioning, which build on the strong resemblance between dopamine neuron activity and the temporal difference prediction error signal in the critic, and between dopamine-dependent long-term synaptic plasticity in...

متن کامل

Application of reinforcement learning to balancing of acrobot

1999

Junichiro Yoshimoto Shin Ishii Masa - aki Sato

The acrobot is a two-link robot, actuated only at the joint between the two links. It is one of dicult tasks in reinforcement learning (RL) to control the acrobot because it has nonlinear dynamics and continuous state and action spaces. In this article, we discuss applying the RL to the task of balancing control of the acrobot. Our RL method has an architecture similar to the actor-critic. The ...

متن کامل

A confidence metric for using neurobiological feedback in actor-critic reinforcement learning based brain-machine interfaces

2014

Noeline W. Prins Justin C. Sanchez Abhishek Prasad

Brain-Machine Interfaces (BMIs) can be used to restore function in people living with paralysis. Current BMIs require extensive calibration that increase the set-up times and external inputs for decoder training that may be difficult to produce in paralyzed individuals. Both these factors have presented challenges in transitioning the technology from research environments to activities of daily...

متن کامل

Two-factor theory, the actor-critic model, and conditioned avoidance.

Journal: :Learning & behavior 2010

Tiago V Maia

Two-factor theory (Mowrer, 1947, 1951, 1956) remains one of the most influential theories of avoidance, but it is at odds with empirical findings that demonstrate sustained avoidance responding in situations in which the theory predicts that the response should extinguish. This article shows that the well-known actor-critic model seamlessly addresses the problems with two-factor theory, while s...

متن کامل