---
type: learning
area: learning
status: learning
date: 2026-05-03
created: 2026-05-03
updated: 1980-01-01
tags:
  - learning
---
- Plan
![[Pasted image 20240922120945 1.png]]
### The Neuron
![[Pasted image 20240922122216 1.png]]

- Artificial neuron
![[Pasted image 20240922123241 1.png]]
Input should standardized or normalized
- Output will be like that
![[Pasted image 20240922123627 1.png]]
By adjusting weight the model became better and better
![[Pasted image 20240922123915 1.png]]
Weights are on synapsis
### What happens in neuron (in General)
1. step 
	 ![[Pasted image 20240922124004 1.png]]
2. step
	Applying an activation function
	 ![[Pasted image 20240922124041 1.png]]
	from this neuron understand if the signal should pass on the signal or no
3. step
	neuron passes the signal

### Activation Function
Is the 2 step in neurons steps
There is 4 types in the activation function

- Threshold function
	![[Pasted image 20240922124702 1.png]]

- Sigmoid 
	advantages is that it's a smooth function
	![[Pasted image 20240922124928 1.png]]
	very useful in case of probabilities

- Rectifier
	![[Pasted image 20240922125009 1.png]]

- Hyperbolic Tangent (tanh)
	![[Pasted image 20240922125053 1.png]]
Very famous combination of functions is applying rectifier one in hidden layers and sigmoid on output

### How does a neural network learn?
- Perceptron
![[Pasted image 20240922155155 1.png]]
 Cost function  = ![[Pasted image 20240922155342 1.png]]

Tells what is the error in the prediction our goal is to minimize it
![[Pasted image 20240922155525 1.png]]
The only thing that we can control in this NN network Wi weights so the process is after getting the output value from NN we compare it to actual value and after based on loss function **C** we update weights until we arrive to minimise C
This process is called back propagation.
### Example:
All of these are the same perceptron
**Epoch** is when we pass the whole dataset one time to train the model
![[Pasted image 20240922161110 1.png]]
### Gradient Descent
Helps to find the best 
Flops = floating operation per second
Best weight = best **Cost** function (most minimized one)
![[Pasted image 20240922162048 1.png]]
Gradient desecent applied in 2d
### The process of applying gradient descent

![[Pasted image 20240922184344 1.png]]
Here you see how **gradient descent** works in a simple graphic representation.
Instead of going through every weight one at a time, and ticking every wrong weight off as you go, **you instead look at the angle of the cost function line**
If the slope is negative, like it is from the highest red dot on the above line, that means you must go downhill from there.
### Stochastic Gradient Descent

![[Pasted image 20240922191222 1.png]]
Applied in not convex function 
and when trying to adjust the weights it we adjust it on after passing every row
It's faster than the normal gradient descent
Stochastic means random
There is another method in between the 2 mini batch gradient descent
### Back propagation

![[Pasted image 20240922193738 1.png]]
is the operation of Optimizing the weights based on the error or cost function C
- Process of Training the ANN with stochastic gradient descent
![[Pasted image 20240922194322 1.png]]

- Number of neurons in output must be decided based on how much variable we want to get in the last example 
	- {0 or 1} -> 1 neuron
	- 3 categorical  -> 3 neurons
- When doing binary classification the loss function should always be **binary_crossentropy** 
- When doing non binary classification the activation function is **soft_max** 