Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
:microscope: Nano size Theano LSTM module
| Date | Stars |
|---|---|
| 2026-07-24 | 304 |
| 2026-07-25 | 304 |
| 2026-07-28 | 304 |
| 2026-07-30 | 304 |
| 2026-08-06 | 304 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
Small Theano LSTM recurrent network module
------------------------------------------
@author: Jonathan Raiman
@date: December 10th 2014
Implements most of the great things that came out
in 2014 concerning recurrent neural networks, and
some good optimizers for these types of networks.
### Key Features
This module contains several Layer types that are useful
for prediction and modeling from sequences:
* A non-recurrent **Layer**, with a connection matrix W, and bias b
* A recurrent **RNN Layer** that takes as input its previous hidden activation and has an initial hidden activation
* A recurrent **LSTM Layer** that takes as input its previous hidden activation and memory cell values, and has initial values for both of those
* An **Embedding** layer that contains an embedding matrix and takes integers as input and returns slices from its embedding matrix (e.g. word vectors)
* A non-recurrent **GatedInput**, with a connection matrix W, and bias b, that multiplies a single scalar to each input (gating jointly multiple inputs)
* Deals with exploding and vanishing gradients with a subgradient optimizer (Adadelta) and element-wise gradient clipping (à la Alex Graves)
This module also contains the **SGD**, **AdaGrad**, and **AdaDelta** gradient descent methods that are constructed using an objective function and a set of theano variables, and returns an `updates` dictionary to pass to a theano function.
### Quick Tutorial
See [a short tutorial for sequence forecasting here](http://nbviewer.ipython.org/github/JonathanRaiman/theano_lstm/blob/master/Tutorial.ipynb).
Or read on for some usage examples.
### Usage
Here is an example of usage with stacked LSTM units, using
Adadelta to optimize, and using a scan operation from Theano (a symbolic loop for backpropagation through time).
dropout = 0.0
model = StackedCells(4, layers=[20, 20], activation=T.tanh, celltype=LSTM)
model.layers[0].in_gate2.activation = lambda x: x
model.layers.append(Layer(20, 2, lambda x: T.nnet.softmax(x)[0]))
# in this example dynamics is a random function that takes our
# output along with the current state and produces an observation
# for t + 1
def step(x, *prev_hiddens):
new_states = stacked_rnn.forward(x, prev_hiddens, dropout)
return [dynamics(x, new_states[-1])] + new_states[:-1]
initial_obs = T.vector()
timesteps = T.iscalar()
result, updates = theano.scan(step,
n_steps=timesteps,
outputs_info=[dict(initial=initial_obs, taps=[-1])] + [dict(initial=layer.initial_hidden_state, taps=[-1]) for layer in model.layers if hasattr(layer, 'initial_hidden_state')])
target = T.vector()
cost = (result[0][:,[0,2]] - target[[0,2]]).norm(L=2) / timesteps
updates, gsums, xsums, lr, max_norm = \
create_optimization_updates(cost, model.params, method='adadelta')
update_fun = theano.function([initial_obs, target, timesteps], cost, updates = updates, allow_input_downcast=True)
predict_fun = theano.function([initial_obs, timesteps], result[0], allow_input_downcast=True)
for example, label in training_set:
c = update_fun(example, label, 10)
### Minibatch usage
Suppose you now have many sequences (of equal length -- we'll generalize this later). Then training can be done in batches:
model = StackedCells(4, layers=[20, 20], activation=T.tanh, celltype=LSTM)
model.layers[0].in_gate2.activation = lambda x: x
model.layers.append(Layer(20, 2, lambda x: T.nnet.softmax(x)[0]))
# in this example dynamics is a function that simulates the behavior of a double
# pendulum and takes our current state and produces an observation
# for t + 1
def dynamics(x, u):
dydx = T.alloc(0.0, 4)
dydx = T.set_subtensor(dydx[0], x[1])
del_ = x[2]-x[0]
den1 = (M1+M2)*L1 - M2*L1*T.cos(del_)*T.cos(del_)
dydx = T.set_subtensor(dydx[1],\n",
( M2*L1 * x[1] * x[1] * T.sin(del_) * T.cos(del_)
+ M2*G * T.sin(x[Excerpt of 9,877 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:cb140177572cbc2c, topic:tutorial, readme:tutorial
matched fp:cb140177572cbc2c, topic:neural-network