> For the complete documentation index, see [llms.txt](https://sejkai.gitbook.io/academic/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://sejkai.gitbook.io/academic/deep-learning/improving-deep-neural-networks/optimization.md).

# Setting up your Optimization Problem

* 在這個篇章裡會講到一些關於優化 cost function 也就是 gradient descent 的方法

## Normalize Inputs

![](https://2991100231-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LhyC30yNfTP1YdCIj83%2F-LsXxlkZOmtOR6BBD8Av%2F-LsXxn5eHhP2B2LKOpy6%2Fnormalize_feature.png?generation=1572544666610247\&alt=media)

* 若某個 feature $$x$$ 的 input range 非常大 (size=1 ... 1000000)
* 那他的 cost function 會像橢圓一樣的形狀
  * 在 gradient descent 時非常緩慢曲折
* 所以我們可以 normalize feature $$x$$
* 使得 cost function 變成像正常的碗狀一樣
  * 適合 implement gradient descent
* Normalize 方法如下
* $$x = \frac{x-\mu}{\sigma}$$
  * 其中的 $$\mu$$ 是平均數 (mean)
    * $$\mu = \frac{1}{m}\sum\_{i=1}^{m}x^{(i)}$$
  * 另一個 $$\sigma$$ 是變異數 (variance)
    * $$\sigma = \sqrt{\frac{1}{m}\sum\_{i=1}^{m}x^{(i)^2}}$$

## Vanishing / Exploding Gradient

* 在 gradient 有可能出現 **vanishing** gradient 或是 **exploding** gradient 的問題
  * 梯度消失、梯度爆炸
* 通常發生在 nn 有非常多 layers 時
* 可被視為阻檔 deep learning 發展的一個原因
* 也就是 weights 將 decrease / increase exponentially
* 假設 activation function 是一個 linear function $$g(z) = z$$
* 假設所有的 $$b^{\[l]} = 0$$
* 如此一來，$$\hat{y}$$ 可以很容易的算出

$$
\hat{y} = W^{\[L]}W^{\[L-1]}\cdots W^{\[2]}W^{\[1]}X
$$

* 若所有 $$W^{\[l]}$$ 的值**大於** 1 時，weights 將會 **increase exponentially**
* 若所有 $$W^{\[l]}$$ 的值**小於** 1 時，weights 將會 **decrease exponentially**
* 對於 backpropogation 的導數一模一樣
* 這會讓訓練難度變得非常高

## Weight Initialization

* Weight initialization 雖然沒辦法完全解決 vanishing / exploding gradient
  * 但可以減緩兩者的發生
* 從 single neuron 看起，若一個 neuron 得到很多個 input，那麼計算 z 等於

$$
z = w\_1x\_1 + w\_2x\_2 +\cdots w\_nx\_n + b
$$

* 因為 n 很大，所以我們勢必要讓每個 $$w\_i$$ 都越小越好
* 為此，我們套用一招 **Xavier initialization** $$\text{Var}(w\_i) = 1/n$$
  * 用 python 寫成

    ```python
    WL = np.random.randn(WL.shape[0], WL.shape[1]) * np.sqrt(1/n)
    ```
  * 其中 n 是输入的神经元个数，即 `WL.shape[1]`
* 若是 activation function 使用 ReLU 時
  * 可以套用另一招 **He initialization** $$\text{Var}(w\_i) = 2/n$$

## Gradient Checking

* Gradient Checking 可以用來檢查 backpropogation 的導數是否正確
* 先利用雙邊誤差方式計算出近似於 slope 的值

![](https://2991100231-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LhyC30yNfTP1YdCIj83%2F-LrjD5sTJNORngfuWK5P%2F-LrjD7B5yoYbmFmnL-jF%2Fgradient_checking_graph.png?generation=1571676523862970\&alt=media)

$$
J'(\theta) = \frac{J(\theta+\epsilon) - J(\theta-\epsilon)}{2\epsilon}
$$

* 當有多個 parameters 時，我們會針對每一個 parameter 都做一次雙邊誤差計算

$$
\begin{aligned}
\text{For each i} &: \\
\&d\theta\_\text{approx}\[i] = \frac{J(\theta\_1,\theta\_2,\cdots,\theta\_i+\epsilon, \cdots) - J(\theta\_1,\theta\_2,\cdots,\theta\_i-\epsilon, \cdots)}{2\epsilon}
\end{aligned}
$$

* 我們希望每一個算出來的值，都可以跟自己計算的一樣
* $$d\theta\_\text{approx}\[i] \approx d\theta\[i] = \frac{\partial J}{\partial \theta\_i}$$
* 我們用以下方式來 check，可以固定比例，不怕數值過小

$$
\frac{\lVert d\theta\_\text{approx} - d\theta \rVert\_2}{\lVert d\theta\_\text{approx}\rVert\_2 + \lVert d\theta \rVert\_2}
$$

* 若結果和 $$\epsilon$$ 近似，表示 backpropogation 做得不錯
* 另外上面計算時用到的為 **Euclidean norm**
  * $$\lVert x\rVert\_2 = \sum\_{i=1}^N\lvert x\_i\rvert^2$$

### Implementation Notes

* Gradient checking 只適用於 debug，training 時應該關掉
* Fail 時，可以去觀察每一個 $$\theta$$ 來找出 bug
* 記得要包含 regularization
* Gradient checking 不適合跟 dropout regularization 一起使用
