Autonomous navigation of stratospheric balloons using reinforcement learning.

Bellemare, Marc G; Candido, Salvatore; Castro, Pablo Samuel; Gong, Jun; Machado, Marlos C; Moitra, Subhodeep; Ponda, Sameera S; Wang, Ziyu

Bellemare, Marc G; Candido, Salvatore; Castro, Pablo Samuel; Gong, Jun; Machado, Marlos C; Moitra, Subhodeep; Ponda, Sameera S; Wang, Ziyu.

Afiliación

Bellemare MG; Brain Team, Google Research, Montreal, Quebec, Canada. bellemare@google.com.
Candido S; Loon, Mountain View, CA, USA. scandido@loon.com.
Castro PS; Brain Team, Google Research, Montreal, Quebec, Canada.
Gong J; Loon, Mountain View, CA, USA.
Machado MC; Brain Team, Google Research, Montreal, Quebec, Canada.
Moitra S; Brain Team, Google Research, Montreal, Quebec, Canada.
Ponda SS; Loon, Mountain View, CA, USA.
Wang Z; Brain Team, Google Research, Toronto, Ontario, Canada.

Nature ; 588(7836): 77-82, 2020 12.

Article en En | MEDLINE | ID: mdl-33268863

ABSTRACT

ABSTRACT

Efficiently navigating a superpressure balloon in the stratosphere1 requires the integration of a multitude of cues, such as wind speed and solar elevation, and the process is complicated by forecast errors and sparse wind measurements. Coupled with the need to make decisions in real time, these factors rule out the use of conventional control techniques2,3. Here we describe the use of reinforcement learning4,5 to create a high-performing flight controller. Our algorithm uses data augmentation6,7 and a self-correcting design to overcome the key technical challenge of reinforcement learning from imperfect data, which has proved to be a major obstacle to its application to physical systems8. We deployed our controller to station Loon superpressure balloons at multiple locations across the globe, including a 39-day controlled experiment over the Pacific Ocean. Analyses show that the controller outperforms Loon's previous algorithm and is robust to the natural diversity in stratospheric winds. These results demonstrate that reinforcement learning is an effective solution to real-world autonomous control problems in which neither conventional methods nor human intervention suffice, offering clues about what may be needed to create artificially intelligent agents that continuously interact with real, dynamic environments.

Texto completo

Añadir a Mi BVS

Imprimir

XML

PubMed Links

Buscar en Google

Texto completo: 1 Colección: 01-internacional Base de datos: MEDLINE Tipo de estudio: Prognostic_studies Idioma: En Revista: Nature Año: 2020 Tipo del documento: Article País de afiliación: Canadá

Texto completo

Añadir a Mi BVS

Imprimir

XML

PubMed Links

Buscar en Google