模型


模型是一个带参数的函数,其中是参数(权重),将输入转化为输出。​

输入

输入就是特征,通常表示为一个向量:。

参数

参数分为和:

  • 权重 :每个输入特征有多重要。
  • 偏置 :一个基础调整量

输出

模型根据输入经过计算后得到的结果,通常表示为。

三要素


模型、策略和算法。

三大范式(三大分类)


监督学习

训练数据既有输入特征,也有标签。

无监督学习

只有输入特征,人工标签。模型需要自己去寻找数据中的潜在规律。

强化学习

过拟合与欠拟合


泛化能力:指模型对没见过的新数据的预测能力。机器学习的终极目标是提高泛化能力。

欠拟合过拟合
概念模型太简单,
没有学到数据的规律。
模型太复杂,
把数据中的噪声和细节也学到了。
训练集表现差极好
测试集表现差差

函数


损失函数

将模型的预测值与真实值映射为一个非负实数,量化模型的误差程度。

激活函数

激活函数是作用于神经元加权输入与偏置之和上的非线性映射函数:
激活函数又分为把饱和激活函数和非饱和激活函数:

  • 饱和激活函数:Sigmoid、Tanh等。
    • 左饱和:
    • 右饱和:
    • 全饱和:同时满足左饱和和右饱和。
  • 非饱和激活函数:不满足饱和的激活函数,如ReLU及其变体等。

Sigmoid函数

{
  "dimension": false,
  "documentSetup": true,
  "title": "Sigmoid",
  "size_x_cm": 10,
  "size_y_cm": 10,
  "show_axis_label": true,
  "axis_label_x": "x",
  "axis_label_y": "y",
  "documentClose": true,
  "showAxis": true,
  "showLargeGrid": true,
  "showSmallGrid": false,
  "gridSize": 5,
  "xmin": "-5.5",
  "xmax": "5.5",
  "ymin": "-0.04187",
  "ymax": "1.04187",
  "axis_style": "middle",
  "functions": [
    {
      "expression": "1/(1+exp(-x))",
      "domain": "-5:5",
      "showLegend": false,
      "fill": false,
      "fillOpacity": 0.2,
      "fillPattern": "solid",
      "tangent": false,
      "dashed": false,
      "tangentPoint": "",
      "extrema": false,
      "color": "black",
      "thickness": "very thin",
      "parametric": false,
      "expressionY": "",
      "name": "Sigmoid"
    }
  ],
  "zmin": "-5",
  "zmax": "5",
  "axis_label_z": "z",
  "rotationX": 30,
  "rotationZ": 45,
  "zoom3D": 1,
  "boxAspect": "true",
  "functions3D": [],
  "majorTickNum": 8,
  "previewSize": 760,
  "annotations": [
    {
      "x": "0",
      "y": "0.5",
      "text": "0.5",
      "color": "black",
      "size": "normal",
      "anchor": "left"
    }
  ],
  "tools": [],
  "coordinateSystem": "cartesian",
  "axis_label_x_polar": "",
  "axis_label_y_polar": "",
  "displayWidth": 700
}

输入范围:
输出范围:
缺点:

  1. 输出不以0为中心,可能导致模型收敛速度慢。
  2. 容易导致梯度消失。

Tanh函数

{
  "dimension": false,
  "documentSetup": true,
  "title": "Tanh",
  "size_x_cm": 10,
  "size_y_cm": 10,
  "show_axis_label": true,
  "axis_label_x": "x",
  "axis_label_y": "y",
  "documentClose": true,
  "showAxis": true,
  "showLargeGrid": false,
  "showSmallGrid": false,
  "gridSize": 5,
  "xmin": "-5.5",
  "xmax": "5.5",
  "ymin": "-1.09988",
  "ymax": "1.09988",
  "axis_style": "middle",
  "functions": [
    {
      "expression": "(exp(x) - exp(-x)) / (exp(x) + exp(-x))",
      "domain": "-5:5",
      "showLegend": false,
      "fill": false,
      "fillOpacity": 0.2,
      "fillPattern": "solid",
      "tangent": false,
      "dashed": false,
      "tangentPoint": "",
      "extrema": false,
      "color": "black",
      "thickness": "thin",
      "parametric": false,
      "expressionY": "",
      "name": "Tanh"
    }
  ],
  "zmin": "-5",
  "zmax": "5",
  "axis_label_z": "z",
  "rotationX": 30,
  "rotationZ": 45,
  "zoom3D": 1,
  "boxAspect": "true",
  "functions3D": [],
  "majorTickNum": 8,
  "previewSize": 760,
  "annotations": [],
  "tools": [],
  "coordinateSystem": "cartesian",
  "axis_label_x_polar": "",
  "axis_label_y_polar": "",
  "displayAlign": "center"
}

输入范围:
输出范围:
优点:

  1. 解决了Sigmoid非零中心问题,收敛更快。
    缺点:
  2. 依然存在梯度消失问题。

ReLU

x, & x \geq 0 \\ 0, & x < 0 \end{cases}$$ ```easy-tikz { "dimension": false, "documentSetup": true, "title": "ReLU", "size_x_cm": 10, "size_y_cm": 10, "show_axis_label": true, "axis_label_x": "x", "axis_label_y": "y", "documentClose": true, "showAxis": true, "showLargeGrid": false, "showSmallGrid": false, "gridSize": 5, "xmin": "-5.5", "xmax": "5.5", "ymin": "-0.245", "ymax": "5.145", "axis_style": "middle", "functions": [ { "expression": "max(0, x)", "domain": "-5:5", "showLegend": false, "fill": false, "fillOpacity": 0.2, "fillPattern": "solid", "tangent": false, "dashed": false, "tangentPoint": "", "extrema": false, "color": "black", "thickness": "thin", "parametric": false, "expressionY": "", "name": "ReLU" } ], "zmin": "-5", "zmax": "5", "axis_label_z": "z", "rotationX": 30, "rotationZ": 45, "zoom3D": 1, "boxAspect": "true", "functions3D": [], "majorTickNum": 8, "previewSize": 760, "annotations": [], "tools": [], "coordinateSystem": "cartesian", "axis_label_x_polar": "", "axis_label_y_polar": "" } ``` **优点**: 1. 不需要计算指数,计算速度快。 2. 缓解梯度消失。 # 符号 --- | | 含义 | | --------------- | ----------------------- | | $x$ | 输入数据(特征) | | $y$ | 真实标签 | | $\hat{y}$ | 模型预测值 | | $N$ / $m$ / $B$ | 样本数量 | | $i$ | 样本编号,$x^{(i)}$表示第$i$个样本 | | $w$ | 权重 | | $b$ | 偏置 | | $\theta$ | 参数的集合,$\theta = (w, b)$ | | $\sigma(\cdot)$ | 激活函数 | | | |