Session 5
效用函数
Session Topic
• 期望货币损益值准则的局限
• 效用函数的定义和公理
• 效用函数的构成
• 风险和效用的关系
• 损失函数、风险函数和贝叶斯风险
期望货币损益值准则的局限
期望货币损益值准则的局限
• 以期望货币损益值为标准的决策方法一般只适用于下列几种
情况:
(1)概率的出现具有明显的客观性值,而且比较稳定;
(2)决策不是解决一次性问题,而是解决多次重复的问题;
(3)决策的结果不会对决策者带来严重的后果。
• 如果不符合这些情况,期望货币损益值准则就不适用,需要
采用其他标准。
• 用期望值作为决策准则的根本条件是,决策有不断反复的可
能。
所谓决策有不断重复的可能,包括下列三层涵义:
第一,决策本身即为重复性决策。
第二,重复的次数要比较多,尤其是当存在对于决策后果有重
大影响的小概率事件时,只有重复次数相当多时才能用期望值
来作为决策标准,因为只有这样其平均后果才接近于后果的期
望值。
例如,要决定是否按月投保火险,而且需要决定的不是投
保一个月(投保一次),而是决定10年内(即120个月)是否投
保,这就重复120次了。但因为失火损失较大,而失火概率又非
常小,比如说仅万分之一,即120个月也不一定会失一次火,所
以其实际平均后果就和期望值相差很大。
假定投保者投保资产为12万元,而保险费规定为万分之二 (保险
费征收率一定比失火概率大,否则保险公司就无盈利可图了),
那么每月应交保险费24元,这是在投保情况下每月的支出。如果
不去投保,则损失的期望值为120000(1/10000)=12元,比投
保的支出小得多。如按期望值标准,则谁也不会去投保了。可是
实际上决策者还是会去投保的,这是因为实际平均损失与其期望
值大不一样,如果这120个月中没有失火,则1元损失也没有,但
万一失火一次,则等于每月平均损失120000/120=1000元,比计
算的损失期望值大80多倍。所以计算出来的损失期望值对决定是
否投保的决策者来说毫无意义,决策者往往会按“不怕一万,只
怕万一”的心里去投保。因为拥有12万元资产的决策者来说,每
月支出24元同其资产额相比几乎等于零,而万一失火却会遭受惨
重损失。
第三,每次决策后果都不会给决策者造成致命的威胁,否则,如
果有此威胁,一旦真的产生此种致命后果,决策者就不可能再作
下一次决策,从而也失去了重复的可能性。这就像投机者把全部
资本孤注一掷一样,一旦失败,资本赔光,下一次也就无法再投
机了。对于有此致命危险的重复性决策,期望值标准的采用也就
受到了限制。
最后,采用期望值标准时,还得假定在不断重复作出相同决策时
其客观条件不变,这一方面包括了个自然状态的概率不变,另一
方面亦包括决策后果函数不变。
Example 1 St. Petersburg paradox
A prime motivator for Bernoulli’s work on the evaluation or risky
ventures was the famous St. Petersburg game. In current terms, a fair
coin is tossed until a head appears. If the first head occurs at the nth
toss, the payoff is 2n$. Suppose you own title to one play of the
game;
that is, you can engage in it without cost. What is the least amount
you would sell your title for? According to the Bernoullis, this least
amount is your equivalent monetary value of the game.
He observed that the expected payoff
(1/2)2 + (1/4)22 + (1/8)23 + … = 1 +1 +1 + …
is infinite, but most people would sell title for a relatively small sum,
and he asked for an explanation of such a flagrant violation of
maximum expected return.
Daniel showed how his theory resolves the issue by providing a unique
solution s to the equation
for any finite w0, where s is the minimum selling price or equivalent
monetary value. Moreover, except for the very rich, a person would
gladly sell title for about $25 or $30. The effect of w0 can be seen
indirectly by estimating your minimum selling price when the payoff
at n is 2n cents instead of 2n dollars and comparing 100 times this
estimate to your answer from the preceding paragraph.
Unlike Bernoulli, Cramer pays little attention to initial wealth, and for
x 0 sets v(x) = . In his terms, the minimum selling price is the
value of s that satisfies
which is a little under $6.
Example 2 A Game illustrating the ‘St. Petersburg paradox’
A casino makes repeated independent tosses of a fair coin until a tail
occurs. A gambler, starting with a stake of $1, is offered the following
wager.
After each toss the gambler will be given two choices. He may
either take away his winnings from the previous tosses of the coin. In
this case the game will end. Alternatively he may use all his winnings
from previous tosses plus his original stake money as a stake for the
next toss of the coin. This stake will be tripled by the casino if a head
is tossed on the next throw. On the other hand, if a tail is thrown the
gambler will lose all his winnings from previous tosses of the coin
together with his original stake money.
Suppose that the gambler is instructed to follow the EMV algorithm
when playing this game. Assume r consecutive heads have been
thrown and denote the gambler’s original stake plus total winnings as
Sr. His expected pay-off for withdrawing from the game is clearly Sr.
However, his expected pay-off for continuing to play is at least (1/2)
3Sr+(1/2)0=(3/2) Sr (the expected pay-off for playing once more). So
under the EMV algorithm the gambler should continue to stake his
winnings until a tail is thrown. But since a tail will be thrown
eventually with probability one, by following the EMV algorithm the
gambler ensures that he will lose his original stake money with
certainty!
Clearly, in the simple game given above, very rational people will not
want to follow the dictates of the EMV algorithm. It is therefore
necessary to modify the EMV algorithm so that ‘optimal decisions’
can be defined sensibly for situations like the one given above. It will
be shown in the next section that such a modification is possible
provided that your client is prepared to commit himself to following
certain rules (or axioms). It also generalizes the EMV approach to
problems when client’s objectives are not only the maximization of
pay-off.
Homework:
A medical laboratory has to test N samples of blood to see which have
traces of a rare disease. The probability any one patient has the disease
is p, and given p, the probability of any group of patients having the
disease is uninfluenced by the existence or otherwise of the disease in
any other disjoint group of patients. Because p is believed to be small it
is suggested that the laboratory combine the blood of d patients into
equal sized pools of n = N/d samples where d is a divisor of N. Each
pool of samples would then be tested to see if it exhibited a trace of the
infection. If no trace were found then the individual samples
comprising the group would be known to be uninfected. If on the other
hand a trace were found in the pooled sample it is then proposed that
each of the n samples comprising that pool be tested individually.
If it costs 1 to test any sample of blood, whether pooled or
unpooled, find the Bayes decision for the optimal size of the groups of
patients for a given value of p.
Answer You are given the space of decisions you are to consider is the
set of divisions of N. All the uncertainty in the experiment exists
because you do not know (d) the number of tests your client will
need to do if he chooses to pool the samples into groups if d samples.
His monetary loss is just L(d)= (d).
The uncertain quantity (d) can be broken down into two
components, being the sum of the number n of tested pools plus the
number of individual patients that subsequently need to be checked. If d
= 1, and he chooses to test samples individually, the second component
of this sum is known to be zero. So the corresponding expected loss
L(1) = N, the number of patients.
Suppose d >1 and you choose to pool the samples in some way. If
denotes the probability that a pool has no trace of diseased blood, then
is the probability that no patient in the pool has the infection. Since
patients have the disease independently it follows by the laws of
probability that
= (1-p)d
So, since he will test n = N/d such pooled samples, the expected
number of pooled samples that need retesting is n(1-). If a pool is
found to have traces of the disease then all members of the pool will be
retested. So the expected number of samples that subsequently need
retesting is
dn(1-) = N{1-(1-p)d}
Adding this number to the chosen number of pools (n = N/d ) gives
the expected number of tests (or equivalently the expected loss in $)
for choosing to use sample pools of size d > 1. Combining these results
gives that
Although is not a linear function of d (as it was in our first
example), given p, can be easily calculated for each divisor d of
N and the decision which minimizes found.
Since is increasing when d > e, you can show that if p
you should choose to test samples individually. On the other hand, if N
is divisible by 3 and p < then it is always optimal to pool samples
in some way. If p < , let y = d-1 + 1- (1-p)d. Then
y = -d-2 – (1-p)dln(1-p) = 0.
We can obtain
Since the right of the equation is constant, this a non-algebra
equation. We may use numerical calculus to approximatively resolve it.
Suppose the resolution is d1, then the optimal d is the divisor of N
which is nearest d1.
效用函数的定义和公理
效用函数的定义和公理
1.效用的概念
决策分析中有两个关键问题:一是对所研究现象的状态的不
确定性进行量化;二是对各种可能出现的后果赋值。一般说来,
状态的不确定性用各种状态出现的概率来描述,而研究出现后果
的价值则要用到效用理论。
所谓效用,就是金钱、物品、劳务或其它事务给人提供的满
足。它是度量一定数量的金钱(或其它事务)在决策者心目中的
价值或者说决策者对待它们的态度的概念。或者说,效用是在有
风险的情况下,决策人对后果的爱好(称为偏好)的量化,可用
一数值表示。 在风险决策中,多用来体现决策者对风险所持有
的态度。
2. 效用函数的定义
定义 展望:设C1,C2,…,Cn表示决策人选择某一行动ai时,
决策问题的全部n个可能的后果;p1,p2,…,pn分别时后果发生
的概率。用P表示所有后果的概率分布,并记为P =(p1, C1; p2,
C2; …; pn, Cn)称为展望。所有展望的集合记作。
定义 在上的效用函数是定义在上的实值函数u :
(1) 它和在上的优先关系 一致,如果对于所有P1, P2 ,
有
P1 P2,当且仅当u(P1) u(P2).
(2) 它在上是线性的,即如果P1, P2 ,而且0 1,则
u(P1+(1-)P2) = u(P1) + (1-)u(P2).
将上述定义推广到一般情况,函数u的线性性可表示为:
如果Pi ,而且i 0, i = 1, 2, …, m, ,则
• 复合展望
由于 P =(p1, C1; p2, C2; …; pn, Cn)
所以 u(P) = u (p1, C1; p2, C2; …; pn, Cn)
记P1=(1, C1; 0, C2; …; 0, Cn) P2=(0, C1; 1, C2; …; 0, Cn) Pn=(0,
C1; 0,
C2; …; 1, Cn),
则P = P1 + P2 + … + Pn
由效用函数在上的现行性质可知,u(P)可表示为
u(P) =
=
=
上式中的u(Ci)为u(1, Ci),即以概率1选择后果Ci的效用。根
据定义,P的效用u(P)就是以概率p1选择后果C1,以概率p2选择后
果C2,……,以概率pn选择后果Cn的期望效用。 因此,如果效
用u存在,而且它和决策人对中的偏好关系一致,即当P1 P2
时,u(P1) u(P2),决策人必将选择一行动使后果的期望效用为
极大。 (举例:带伞问题)
理性行为公理:
公理1 连通性(或成对可比性):如果P1, P2 ,则或者P1
P2,或者P1 P2,或者P1 P2 。
公理2 传递性:如果P1, P2, P3,而且P1 P2,P2 P3,则
必有P1 P3 。
公理3 替代性:如果P1, P2 和Q,而且0<p<1,则
P1 P2 当且仅当 pP1 + (1-p)Q pP2 + (1-p)Q .
公理4 连续性(连续性或称偏好有界性):
如果P1, P2, P3,而且P1 P2 P3,则存在数p和q,
0<p<1和0<q<1,使
pP1 + (1-p)P3 P2 qP1 + (1-q)P3
公理1和公理2合称为次序性公理。符合次序性公理的集合称
为全(弱)序集,集中的元素可以按偏好关系排列优先次序,表
示的是决策者对行动的偏爱程度的比较。这两条公理是说,对行
动的偏爱是可以比较的。公理3是说,偏好关系中的两个有序后
果在各有相同比例 (1-p) 被相等量(1-p)Q 替代后,优先关系不变。
公理4意味着没有一种后果无限好,也没有一种后果无限坏。公
理4还可以用如下方式表示:若P1 P2 P3,则必有0 1,
使 P2 P1 + (1-)P3 (称为效用值)。这条公理告诉我们:对于
较好行动后果C1,较差行动后果C3及中间行动后果C2,总可以调
整P的大小,使得复合行动“以概率P获C3,以概率1-p获C1”与
行动后果C2相比较,决策者同样偏爱。这条公理有时不易被人们
接受,特别是当较差行动导致严重后果时。
定理 在上的优先关系 如满足公理1至公理4,则在上存
在一效用函数u,它和 一致。此外,u经过正线性变换,仍然是
和 一致的效用。
Theorem If u() is a utility function on X, then w() = u()+
( >0) is also a utility function representing the same preferences.
Conversely, if u() and w() are two utility functions on X representing
the same preferences, then there exist >0 and such that w() =
u()+.
Proof. The first part of the theorem is one part of the theorem . To
prove the converse implication suppose that u() and w() are two utility
functions on X representing the same preferences. Suppose also that
w() u()+ (5. 1)
We will obtain a contradiction. For any point xi X define a point (ui,
wi) in the plane, where ui = u(xi) and wi = w(xi). If ( ) holds, xi
X (i = 1, 2, 3) such that the points
(u1, w1), (u2, w2) and (u3,, w3)
are not colinear. Without loss of generality we may assume that
u1<u2<u3 . Let p = (u2-u1)/(u3-u1) and consider the lottery x3px1,
where x3px1 = <(1-p), x1; 0, x2; p, x3>. In terms of the function
u(),x3px1 has expected utility.
= pu(x3) + (1-p) u(x1)
=
= u2
So x2 x3px1 . But simple geometry shows that the assumption of non
-
collinearity implies
W2 pw3 + (1-p) w1 .
See Fig. So in terms of the utility function w(),x2 x3px1 .
Hence the assumption that u() and w() represent the same
preferences is contradicted. Therefore ( ) cannot hold and we have
w() = u()+ . That >0 holds clearly.
uu1 u2 u3
w
w3
w2
w1
pw3+(1-p)w1
确定当量是指以下两种情况等价:一种情况是决策人得到一确定
的后果C1,另一种情况是决策人得到一抽奖的机会(记作C) ,他
以概率p得到后果C2,以概率1-p得到后果C3,即(p, C2; (1-p), C3
)。如果决策人认为这两种情况对他是等价的,则确定的后果C1
称为抽奖(p, C2; (1-p), C3)的确定当量。
u(C1) = E[u(C)]
公理5(抽奖的性质) 一抽奖的所有奖金都增加一金额 ,将使
此抽奖的确定当量增加 。
定理 如果在抽奖集上的优先关系 适合公理1至公理5,则
在后果集X上的效用函数为线性函数或指数函数。
证明:假设抽奖的奖金(后果)为一连续变量,记作x,奖金集为
X,X上的概率密度为f(x). 此时抽奖P的期望效用为
根据公理5和确定当量的定义可知,
()
其中 是抽奖的确定当量。假设为一连续变量,在后果集X上
的效用函数u(x+)对有直到二阶的连续导数。将上式两边对
位分两次,得
和
将以上两式左右两边分别相除,并令=0,得
如果()式对于各种不同的密度函数都成立,则满足该式的效
用函数应符合以下条件:
()
对所有的xX
选择此比值为一常数-r,即
上式积分,得
lnu(x) = -rx + k0
所以 u(x) = k1e
-rx
如r = 0,则有u(x) = k1x+k2
如r 0,则 u(x) = k2e
-rx+k3
Example 3 The St Petersburg paradox revisited
In Example 2 assume now that the gambler has a utility function on
reward r of the form
u(r) = r(e + r)-1 r > 0 ()
The parameter will reflect the gambler’s propensity to take risks.
The larger the value of , the more he prefers speculative gains of
large amounts to the certainty of gaining small amounts. Regardless of
, u(r) is concave and takes values between 0 and 1.
The decision-maker’s expected utility associated with decision dn of
terminating the game after the nth toss is
= (1/2)n-13n-1(e + 3n-1)-1 = ()n-1(e + 3n-1)-1
It is easily checked that a function ex [e + ex]-1 is maximized when
x = -1[log - log() + ] if
It follows that ()y-1(e + 3y-1)-1 is maximized when
y = -1[log - log() + ] + 1
where = = log3
is therefore maximized at either the largest non-negative
integer value n less than y or the smallest non-negative integer greater
than y, where y is defined above. Notice in particular that y increases
(linearly) with . So the more ready the gambler is to risk large
speculative gains, the longer he should play the game, as we would
expect. But note that any gambler with a utility function given in
equation () will choose to terminate the game at some stage –
unlike the EMV decision-maker. So then the supposed St Petersburg
‘paradox’ has disappeared.
基数效用:以上所定义的效用是决策人在有风险的情况下对后果
的偏好的量化,其中含有决策人对一不确定事件可能冒的风险的
态度,它反映了决策人的偏好强度,这种效用称为基数效用。
序数效用:定义一效用表示决策人对各种确定事件的后果的偏好
次序。对于这类事件,决策人无需承担任何风险。这样定义的效
用和基数效用不同,称为序数效用。
定义 令X为所有确定事件的后果x的集,在X上的效用函数称
为序数效用函数,它是定义在X上的实值函数u,有
u(x1) u(x2),当且仅当 x1 x2。
如在X上的优先关系 满足以下三条公理,则在X上的序数效
用存在:
公理1 连通性、公理2 传递性、公理3 连续性,即对于任何确定
的后果x,它的劣势集和它的优势集都是闭集。
定理 如果在X上的优先关系 满足公理1至公理3
,
则在X上存在效用u,它和 一致。此外,u经过保序
变换仍然和 一致。
效用函数的构成
1. 离散型效用的测定
效用的大小可以用概率的形式来表示,效用值介于0、1之
间。效用的测定方法很多,最常用的是Von Neumann和
Morgenstern于1944年共同提出的,称之为效用标准测定法。
例如,对决策者而言,他最大的愿望是收益100元,最小
的收益为0元。将收益为100元的方案的效用指定为1,收益为0
元的方案的效用值定为0,且分别记作u(100)=1, u(0)=0。
如果决策者确定“u(100)=1, u(0)=0”,那么记“收益a元
的行
动方案”为a,0<a<100 。下面确定u(a)。
显然,对决策者愈有利的方案,效用值愈大,因此u(a)应
满足0<u(a)<1。
效用函数的构成
现以a=50为例,确定u(50)的值。首先可向决策者提出如
下问题:“现有两个行动方案a1和a2,如果采取行动方案a1,能
以的概率获0元和的概率获100元的收益;如果采取行动
方案a2,肯定能得到50元的收益。请问,你愿意采取哪个行动
方案?”
如果决策者选择a2,我们再提第二个问题:如果行动方案a1
的两个可能结果的概率发生变化,即以的概率获得0元,以
的概率获100元,行动方案a2不变,决策者如何选择?假如此次
决策者的选择为a1,我们再适当升高方案a1取0的概率,降低获
100的概率,再次提问。如此继续下去,直到决策者认为采取行
动方案a1与采取行动方案a2对他来说是同样时为止。此时,行动
方案a1与a2在决策者心目中的地位平等,即称为等效行动。
(说明效用的主观性)
假设此时的方案a1获0元的概率为,获100元的概率为,
因为在决策者看来,行动方案a1和a2的效用相等,故
(0)+(100)=u(50)
u(50)=
提出上述问题时,行动方案a1的两个可能结果不一定用0元和
100元的效用值比较,也可以用任意两个货币值,只要它们的
效用已经得出,并且预测点介于二者之间。
对于货币效用,如果知道了效用值,也可以测出货币值。比
如货币的效用值为,我们来测出相应的货币值。可以这样
提问:
“现有两个行动方案a1和a2,如果采取行动方案a1,能以的
概率获0元和的概率获100元的收益;如果采取行动方案a2,
肯定能得到40元的收益。请问,你愿意采取哪个行动方案?”
假设决策者的回答是选择行动方案a2,则可以把40元降低
一些,如将为25元,再向他提同样的问题。如果它这时选择
a1,我们就把25元适当提高,再要求它比较两个方案,直到决
策者认为两个方案等效为止。假设此时a2肯定获得的收益为30
元,则有:
u(30) = (0)+(100)=
也可以调整数之后继续提问。这样就可以得出任一手一值
介于0和100之间的效用值。将这些值[c, u(c)]以平滑的曲线连
接
起来,便得到决策者的效用曲线。
效用的概念不仅适用于货币,而且适用于非货币事物,即
运用标准测定法,亦可测出非货币事物的效用。比如,对某决
策者来说,又A、B、C、D、E五件事情,我们将测定它们对
决策者的效用。具体做法是:首先要确定决策者最喜爱和最不
喜爱那件事情。假设他最喜爱的是A,最不喜爱的是E,则令
u(A)=1, u(E)=0,如果要测定B的效用,可以这样提问:
“有行动方案a1和a2,方案a1可能以p的概率使你获得A和以(1-p)
的概率获得E;方案a2有1的概率获得B。你认为当p为何值时,
方案a1与a2等效?”
决策者在回答p值后,便可以同样的计算方法得出u(B)。同样
也可求出u(C)和u(D)。
2. 连续型效用的测定
P43 例
效用函数的确定
确定效用函数是一个相当复杂的过程,它至少需要解决三方面
的问题:其一,测定出集合X中一些离散点的效用;其二,在不
同条件下,建立描述主体偏好特性的数学模型;其三,效用函数
的导出和确定。
用标准测定法来构造效用函数,其实质是采用问答的方式,对
决策者作一些心理测验,通过回答以了解决策者对随机事件与确
定型事件在效用值上的等价关系,通过货币值及所求的对应效用
值的坐标关系,就可以得出足够多的坐标点,而后把这些点以平
滑的曲线连接起来,便得到决策者的效用曲线。但此法比较麻
烦,因为要对决策者作反复的提问。为了给出两种比较简便的方
法,以减少对决策者的提问次数,需要先介绍风险与效用的关
系。
风险和效用的关系
风险和效用的关系
自从1944年Neumann和Morgenstern将期望效用假设公
理化后,经济学家们就迅速将其应用于诸如投资选择、
保险等经济领域中。
• 对风险的态度
(1)风险厌恶
(2)风险中立
(3)风险喜好
• 例:
确定性后果u(2500)=1,u(0)=0,抽奖a2后果的期望值为
1250,Eu(a2)=
风险和效用的关系
• 风险酬金
在风险厌恶的情况下: s=E(x)-k
s为确定当量,E(x)为后果的期望值,k为风险酬金
(1)k>0,风险厌恶
(2)k=0,风险中性
(3)k<0,风险喜好
可测价值函数
• 后果的偏好强度
0-1000与1000-1500等价,效用函数?
对确定性后果的偏好强度需要测量
• 序数价值函数
设 是定义在方案集A上的决策人的弱序,
若
A上的实值函数v满足
v(a)≥v(b)<=>a b
可测价值函数
• 为了量化决策人对确定性后果的偏好强度
• 可测价值函数
在后果X上的实值函数v,对w,x,y,z属于X,有
(1)(w→x) (y→z)<=>v(w)-v(x) ≥ v(y)-v(z)
(2)v对正线性变换是唯一确定的
(w→x)表示决策人对w和x的偏好强度之差
• 可测价值函数示意图
相对风险态度
• 决策人的真实的风险态度被称为“相对风险态度”
• 假设效用函数u和可测价值函数v在X上单调递增且二次
连续可微
记
表示决策在x处的风险侧度
(1)r(x)>0,风险厌恶
(2)r(x)=0,风险中性
(3)r(x)<0,风险喜好
偏好强度的局部侧度
• 可测价值可以反映决策人在x处的偏好强度,
即边际价值m(x)
(1)m(x)>0,v在x处下凹,边际价值递减
(2)m(x)=0,v在x处线形,边际价值不变
(3)m(x)<0,v在x处凸,边际价值递增
真正的风险态度
• 决策人的真实风险态度
效用函数的常见形式
• 定常风险厌恶效用函数
u(x)>0, u(x)<0, 相应的效用函数为
凹的,r(x)=c>0 。
效用函数具有如下形式:
u(x)=a-be-cx , c>0, b>0
2) 递增风险厌恶效用函数
递增风险厌恶是指当主体随着其财产水平x的增加,他对某一类
特定范围的决策就愈加回避风险的态度。这时r(x)>0,u(x)>0,
u(x)<0, 相应的效用函数为凹的。
当选取 ,a, b>0, 0<x<b
这时r(x)>0,u(x)>0, u(x)<0, 相应的效用函数为凹的。
u(x)=c-d(b-x)a+1
当a=1时,效用函数具有二次函数的形式,它是递增风险厌恶中
较
为实用的一种效用函数。
3) 递减风险厌恶效用函数
当主体的风险态度是递减风险厌恶时,这时r(x)>0,u(x)>0,
u(x)<0, 相应的效用函数为凹的,但与(2)不同的是,u(x)的凹
性
随着x的增加而逐渐减弱。
当选取 r(x)=a/x , a>0 ,根据参数a的不同,可以导出以下三种
形
式的基本递减风险厌恶效用函数。
• 0<a<1 一般形式为u(x)=bxq+d
• a=1 u(x)=c1lnx+c2
• a>1 一般形式为u(x)=-bx+d
损失函数、风险函数和
贝叶斯风险
损失函数、风险函数和贝叶斯风险
损失函数记作l(, a),它表示一决策问题当状态为,决策人
的行动为a时,所产生的后果使决策人遭受的损失。由于损失
函数可能为负值,因此它也能反映决策人获得的收益。后果的
效用越大,损失越小。故用效用函数去定义损失函数的一种简
单办法,是令
l(, a) = - u(, a)
为了使损失函数非负,可以定义为
由于损失函数经过任何正线性变换仍然是同一优先关系的效用函
数,因此以上两种形式的损失函数都会得到同样的分析结果。
对于给定的,观察的结果X是一随机变量,用F(x)记X的
条
件分布函数,用f(x)记X的条件密度函数,用 记随机变量
的
样本空间。
决策规则
所谓决策法则,就是由所有可能信息值的集合到所有可能行动的
集合的一个映射。换句话说,决策规则是这样一个规则,按照这
个规则,对于每一个信息值X均有唯一确定的可行行动a=(x)与
之
对应。
设给定一个决策规则(x),在任一状态下,当信息值X确定后,
它所对应的行动(x)也就确定了,从而(x)的损失值为l(, (x)
),
它也是一随机变量。当给定,l(, (x) )对X的期望值称为风险
函
数,并记为R(, ),
R(, ) = E
x l(, (x) )
若X为离散随机变量,则
由于决策认事先不知道真实状态,他只能对随机状态的先验密
度()作出主观估计。所以进行决策分析时,还需要将损失函数
R(, (x) )对取期望值,即
如为连续随机变量,则
如为离散随机变量,则
r(, )称为决策规则相对于的贝叶斯风险。
对于固定的决策规则(x),其贝叶斯风险为一常数,它反映出利
用这一决策规则决策的平均损失。
三种标准“损失函数”:
• 平方损失:
更一般的平方损失是加权平方损失,其形式是
• 线性损失
• “0-1”损失
例. 有两类盒子:甲类盒子只有一个,其中装有80个红球,20
个白球;乙类盒子共有三个,每个盒子均装有20个红球,80个白
球。四个盒子外表一样,内容不知。今从中任取一盒,请你猜它
是哪类的。如果猜中,付你1元钱;如果未猜中,不付你钱。那
么你怎样猜法?
如果从这个盒子中任意抽取N个球(回置地),让你观察,你
如何根据这N个球的性质来选择自己的行动?
当容量为1或2的抽样时,求各决策规则的风险函数和贝叶斯风
险,并分别指出最佳决策规则。
解:令表示所取出的这个盒子中红球所占的比例。显然,
只能取两个值:若这个盒子是甲类的, = 1= ;若这个盒子
是乙类的, = 1= 。用a1、a2分别表示猜这个盒子是甲类的和
猜它是乙类的这两个行动方案。显然收益矩阵如表5. 1所示。
表5. 1 猜盒问题的收益矩阵
1 2
1/4 3/4
a1 1 0
a2 0 1
假设N=1,即从你猜的那个盒子中取出1个球来观察。规定:对于
红球,x=1;对于白球,x=0,其抽样分布如表 所示:
表5. 2 N=1时猜盒问题的抽样分布矩阵
P() P(x=0/) P(x=1/
)
1 1/4
2 3/4
P(x=0/)表示从甲类盒子中抽取1球是白球的概率,显然它等
于 。对另外三个概率可作类似理解。
利用先验分布和抽样分布计算后验分布:
P(x=0)=1/4+3/4=
P(1/x=0)=
P(2/x=0)=
P(x=1)=1/4+3/4=
P(1/x=1)=
P(2/x=1)=
表5. 3 N=1时猜盒问题的后验分布矩阵
X P(x) P(1/x) P(2
/x)
0 1/13 12/13
1 4/7 3/7
样本容量N=2时,
表5. 4 N=2时猜盒问题的抽样分布矩阵
P()
P(x=0/)
P(x=1/) P(x=2/)
1 1/4
2 3/4
表 N=2时猜盒问题的后验分布矩阵
X P(x) P(1/x) P(2/x)
0 1/49 48/49
1 1/4 3/4
2 16/19 3/19
本例中,如果样本容量为1,由于所有可能的抽样结果有2个,可
行行动也有2个,故决策规则共有22 = 4个:
1(x)=a1
2(x)=
3(x)=
4(x)=a2
如果样本容量为2,那么抽样结果有3种可能,可行行动海时2个,
因此决策规则共有23=8个。
一般地,对于有S个可行行动的决策问题,若补充信息值有n个,
则决策规则共有Sn个。
对于决策法则1(x),无论是x=0或x=1,都有
l(1, 1(x)) = l(1, a1) = 0
l(2, 1(x)) = l(2, a1) =1
于是
R(1, 1(x)) = l(1, 1(x)) = 0
R(2, 1(x)) = l(1, 1(x)) =1
即1(x)的风险函数为
其贝叶斯风险:
r(, 1) = R(1, 1)P(1)+ R(2, 1)P(2) =
对于决策规则2(x),
l(1, 2(0)) = l(1, a1) = 0
l(1, 1(1)) = l(1, a2) = 1
于是
R(1, 2(x)) = R(1, 2(0))P(x=0/1)+ R(1,
2(1))P(x=1/1)
= 1+0=
同样地,
l(2, 2(0)) = l(2, a1) = 1
l(2, 2(1)) = l(2, a2) = 0
R(2, 2(x)) = R(2, 2(0))P(x=0/2)+ R(2,
2(1))P(x=1/2)
= 1+0=
因此,2(x)的风险函数为R(, 2) =
其贝叶斯风险为r(, 2)=
类似地,可以求出3和 4的贝叶斯风险分别为r(, 3)=,
r(, 4)=
最佳决策规则为3
贝叶斯准则:贝叶斯风险最小的决策规则为最佳决策规则。
最小最大风险函数值准则:对于每个决策规则(x),找出它的风
险
函数在所有可能状态下的最大值,这些最大风险函数值中最小者
所对应的决策规则,就是最小最大值准则下的最佳决策规则。
注意:这里所将的决策规则与前面的决策准则不同。决策准则,
是判别诸行动间优劣关系的标准,他所指明的是什么叫一个行动
方案优于另一个行动方案,据此可以选出最优行动方案来;而决
策法则是信息值与所采取的行动的对应关系,它所指明的是如何
根据信息值选择行动方案。