Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Deep Reinforcement Learning For Sequence to Sequence Models
| Date | Stars |
|---|---|
| 2026-07-31 | 767 |
| 2026-08-01 | 767 |
| 2026-08-02 | 767 |
| 2026-08-06 | 767 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
********************
RLSeq2Seq
********************
.. image:: https://img.shields.io/badge/contributions-welcome-brightgreen.svg?style=flat
:target: https://github.com/yaserkl/RLSeq2Seq/pulls
.. image:: https://img.shields.io/badge/Made%20with-Python-1f425f.svg
:target: https://www.python.org/
.. image:: https://img.shields.io/pypi/l/ansicolortags.svg
:target: https://github.com/yaserkl/RLSeq2Seq/blob/master/LICENSE.txt
.. image:: https://img.shields.io/github/contributors/Naereen/StrapDown.js.svg
:target: https://github.com/yaserkl/RLSeq2Seq/graphs/contributors
.. image:: https://img.shields.io/github/issues/Naereen/StrapDown.js.svg
:target: https://github.com/yaserkl/RLSeq2Seq/issues
.. image:: https://img.shields.io/badge/arXiv-1805.09461-red.svg?style=flat
:target: https://arxiv.org/abs/1805.09461
NOTE: This code is no longer actively maintained.
This repository contains the code developed in TensorFlow_ for the following paper:
| `Deep Reinforcement Learning For Sequence to Sequence Models`_,
| by: `Yaser Keneshloo`_, `Tian Shi`_, `Naren Ramakrishnan`_, and `Chandan K. Reddy`_
.. _Deep Reinforcement Learning For Sequence to Sequence Models: https://arxiv.org/abs/1805.09461
.. _TensorFlow: https://www.tensorflow.org/
.. _Yaser Keneshloo: https://github.com/yaserkl
.. _Tian Shi: http://life-tp.com/Tian_Shi/
.. _Chandan K. Reddy: http://people.cs.vt.edu/~reddy/
.. _Naren Ramakrishnan: http://people.cs.vt.edu/naren/
If you used this code, please kindly consider citing the following paper:
.. code:: shell
@article{keneshloo2018deep,
title={Deep Reinforcement Learning For Sequence to Sequence Models},
author={Keneshloo, Yaser and Shi, Tian and Ramakrishnan, Naren and Reddy, Chandan K.},
journal={arXiv preprint arXiv:1805.09461},
year={2018}
}
#################
Table of Contents
#################
.. contents::
:local:
:depth: 3
.. Chapter 1 Title
.. ===============
.. Section 1.1 Title
.. -----------------
.. Subsection 1.1.1 Title
.. ~~~~~~~~~~~~~~~~~~~~~~
.. image:: docs/_img/rlseq.png
:target: docs/_img/rlseq.png
============
Motivation
============
In recent years, sequence-to-sequence (seq2seq) models are used in a variety of tasks from machine translation, headline generation, text summarization, speech to text, to image caption generation. The underlying framework of all these models are usually a deep neural network which contains an encoder and decoder. The encoder processes the input data and a decoder receives the output of the encoder and generates the final output. Although simply using an encoder/decoder model would, most of the time, produce better result than traditional methods on the above-mentioned tasks, researchers proposed additional improvements over these sequence to sequence models, like using an attention-based model over the input, pointer-generation models, and self-attention models. However, all these seq2seq models suffer from two common problems: 1) exposure bias and 2) inconsistency between train/test measurement. Recently a completely fresh point of view emerged in solving these two problems in seq2seq models by using methods in Reinforcement Learning (RL). In these new researches, we try to look at the seq2seq problems from the RL point of view and we try to come up with a formulation that could combine the power of RL methods in decision-making and sequence to sequence models in remembering long memories. In this paper, we will summarize some of the most recent frameworks that combines concepts from RL world to the deep neural network area and explain how these two areas could benefit from each other in solving complex seq2seq tasks. In the end, we will provide insights on some of the problems of the current existing models and how we can improve
them with better RL models. We also provide the source code for implementing most of the models that will be discussed in this paper on the complex task of Excerpt of 30,301 characters
Read on GitHub140
Sina Torfi · Meta · United States
36
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:5eee2526abf956f7, topic:reinforcement-learning, desc:reinforcement learning
matched fp:5eee2526abf956f7, topic:nlp