2020-09-22 21:56 已编辑门头沟学院产品经理

关注

【数据分析学习笔记day12】 Pandas的索引操作+索引对象Index+Series索引+DataFrame索引+高级索引：标签、位置和混合+1. Series和DataFrame中的索引都是In

文章目录

Pandas的索引操作

Pandas的索引操作

索引对象Index

1. Series和DataFrame中的索引都是Index对象

示例代码：

print(type(ser_obj.index))
print(type(df_obj2.index))

print(df_obj2.index)

运行结果：

<class 'pandas.indexes.range.RangeIndex'>
<class 'pandas.indexes.numeric.Int64Index'>
Int64Index([0, 1, 2, 3], dtype='int64')

2. 索引对象不可变，保证了数据的安全

示例代码：

# 索引对象不可变
df_obj2.index[0] = 2

运行结果：

---------------------------------------------------------------------------
TypeError                                 Traceback (most recent call last)
<ipython-input-23-7f40a356d7d1> in <module>()
      1 # 索引对象不可变
----> 2 df_obj2.index[0] = 2

/Users/Power/anaconda/lib/python3.6/site-packages/pandas/indexes/base.py in __setitem__(self, key, value)
   1402 
   1403     def __setitem__(self, key, value):
-> 1404         raise TypeError("Index does not support mutable operations")
   1405 
   1406     def __getitem__(self, key):

TypeError: Index does not support mutable operations

常见的Index种类

Index，索引
Int64Index，整数索引
MultiIndex，层级索引
DatetimeIndex，时间戳类型

Series索引

1. index 指定行索引名

示例代码：

ser_obj = pd.Series(range(5), index = ['a', 'b', 'c', 'd', 'e'])
print(ser_obj.head())

运行结果：

a    0
b    1
c    2
d    3
e    4
dtype: int64

2. 行索引

ser_obj[‘label’], ser_obj[pos]

示例代码：

# 行索引
print(ser_obj['b'])
print(ser_obj[2])

运行结果：

1
2

3. 切片索引

ser_obj[2:4], ser_obj[‘label1’: ’label3’]

注意，按索引名切片操作时，是包含终止索引的。

示例代码：

# 切片索引
print(ser_obj[1:3])
print(ser_obj['b':'d'])

运行结果：

b    1
c    2
dtype: int64
b    1
c    2
d    3
dtype: int64

4. 不连续索引

ser_obj[[‘label1’, ’label2’, ‘label3’]]

示例代码：

# 不连续索引
print(ser_obj[[0, 2, 4]])
print(ser_obj[['a', 'e']])

运行结果：

a    0
c    2
e    4
dtype: int64
a    0
e    4
dtype: int64

5. 布尔索引

示例代码：

# 布尔索引
ser_bool = ser_obj > 2
print(ser_bool)
print(ser_obj[ser_bool])

print(ser_obj[ser_obj > 2])

运行结果：

a    False
b    False
c    False
d     True
e     True
dtype: bool
d    3
e    4
dtype: int64
d    3
e    4
dtype: int64

DataFrame索引

1. columns 指定列索引名

示例代码：

import numpy as np

df_obj = pd.DataFrame(np.random.randn(5,4), columns = ['a', 'b', 'c', 'd'])
print(df_obj.head())

运行结果：

          a         b         c         d
0 -0.241678  0.621589  0.843546 -0.383105
1 -0.526918 -0.485325  1.124420 -0.653144
2 -1.074163  0.939324 -0.309822 -0.209149
3 -0.716816  1.844654 -2.123637 -1.323484
4  0.368212 -0.910324  0.064703  0.486016

[外链图片转存失败,源站可能有防盗链机制,建议将图片保存下来直接上传(img-k9JqJxrj-1579951879137)(…/images/DataFrameIndex.png)]

2. 列索引

df_obj[[‘label’]]

示例代码：

# 列索引
print(df_obj['a']) # 返回Series类型
print(df_obj[[0]]) # 返回DataFrame类型
print(type(df_obj[[0]])) # 返回DataFrame类型

运行结果：

0   -0.241678
1   -0.526918
2   -1.074163
3   -0.716816
4    0.368212
Name: a, dtype: float64
<class 'pandas.core.frame.DataFrame'>

3. 不连续索引

df_obj[[‘label1’, ‘label2’]]

示例代码：

# 不连续索引
print(df_obj[['a','c']])
print(df_obj[[1, 3]])

运行结果：

          a         c
0 -0.241678  0.843546
1 -0.526918  1.124420
2 -1.074163 -0.309822
3 -0.716816 -2.123637
4  0.368212  0.064703
          b         d
0  0.621589 -0.383105
1 -0.485325 -0.653144
2  0.939324 -0.209149
3  1.844654 -1.323484
4 -0.910324  0.486016

高级索引：标签、位置和混合

Pandas的高级索引有3种

1. loc 标签索引

DataFrame 不能直接切片，可以通过loc来做切片

loc是基于标签名的索引，也就是我们自定义的索引名

示例代码：

# 标签索引 loc
# Series
print(ser_obj['b':'d'])
print(ser_obj.loc['b':'d'])

# DataFrame
print(df_obj['a'])

# 第一个参数索引行，第二个参数是列
print(df_obj.loc[0:2, 'a'])

运行结果：

b    1
c    2
d    3
dtype: int64
b    1
c    2
d    3
dtype: int64

0   -0.241678
1   -0.526918
2   -1.074163
3   -0.716816
4    0.368212
Name: a, dtype: float64
0   -0.241678
1   -0.526918
2   -1.074163
Name: a, dtype: float64

2. iloc 位置索引

作用和loc一样，不过是基于索引编号来索引

示例代码：

# 整型位置索引 iloc
# Series
print(ser_obj[1:3])
print(ser_obj.iloc[1:3])

# DataFrame
print(df_obj.iloc[0:2, 0]) # 注意和df_obj.loc[0:2, 'a']的区别

运行结果：

b    1
c    2
dtype: int64
b    1
c    2
dtype: int64

0   -0.241678
1   -0.526918
Name: a, dtype: float64

3. ix 标签与位置混合索引

ix是以上二者的综合，既可以使用索引编号，又可以使用自定义索引，要视情况不同来使用，

如果索引既有数字又有英文，那么这种方式是不建议使用的，容易导致定位的混乱。

示例代码：

# 混合索引 ix
# Series
print(ser_obj.ix[1:3])
print(ser_obj.ix['b':'c'])

# DataFrame
print(df_obj.loc[0:2, 'a'])
print(df_obj.ix[0:2, 0])

运行结果：

b    1
c    2
dtype: int64
b    1
c    2
dtype: int64

0   -0.241678
1   -0.526918
2   -1.074163
Name: a, dtype: float64

注意

DataFrame索引操作，可将其看作ndarray的索引操作

标签的切片索引是包含末尾位置的

全部评论

推荐最新楼层

不愿透露姓名的神秘牛友

11-19 23:05

26届日常实习简历求各位大佬拷打指点

bg:哈工大本硕非科班，无实习，想找Java后端的日常实习；项目方面：一个苍穹外卖，一个黑马商城的微服务，一个自己做着玩的很简单的项目，3个不相关的实验室项目；在牛客、Boss、实习僧上投了一些简历，根本没多少人看，看了的也都是要完简历就没消息了

逍遥生777：你找java的后端开发，那和java无关的项目就不用写了，剩余的项目写详细点

投递牛客等公司10个岗位 > 简历中的项目经历要怎么写实习，投递多份简历没人回复怎么办

点赞评论收藏

昨天 12:32

长春理工大学金融分析师

重生之我变成了小学生

家人们！大离谱事件发生了！早上我被我妈叫醒，我就想，我都是一个20多岁的成年人了，怎么早上还叫我起床！所以我就没理然后我的屁股遭到了一记重击！等等……这感觉怎么似曾相识……难道是……难道是……我猛然睁开眼睛！日历上显示现在是，，2012年！！我穿越了！！我呆呆地坐在床上，看着我现在幼小的身体，我妈拎着笤帚站在床头，看我愣着出神，超用力地推了一下我的头，顺势我倒在了床上……“你看看这都几点了还不起？不想上学啦！不想上学早说，跟着你爸干活去！别给我浪费钱！”说罢她“嘭”地把门关上离开了难道说……我真的回到了小学？我迅速爬起来，把自己浑身上下摸了一遍，给了自己两耳刮子，啧……真疼啊我，真的回到小学了...

OfferLetters：我可不想再回去从小学开始重新上学，每天起的比鸡早，睡得比狗晚，高中三年透支精神，透支身体，比杀了我还难受

非技术求职现状

点赞评论收藏

10-05 23:02

东北大学 Java

很难想象仅仅过去了4年找工作的难度就大不一样

我说句实话啊：那时候看三个月培训班视频，随便做个项目背点八股，都能说3 40w是侮辱价

点赞评论收藏

11-09 17:30

门头沟学院 Java

抽象华为

TYUT太摆金星：我也是，好几个华为的社招找我了

点赞评论收藏

11-19 16:05

中山大学业务管理

半夜无聊去查了一下社保后

惊喜地发现公司还给交了一金（因为面试的时候说只有五险），然后库库一顿查，想看看交的哪个档位 结果天还是塌了… 因为面试的时候薪资报低了，一顿猛算后发现还是比招聘岗位上最低的薪资低… ok fine，i’m a cheap man i know

点赞评论收藏

点赞收藏评论

全站热榜

正在热议

# 选完offer后，你后悔学本专业吗 #

20890次浏览 149人参与

# 阿里云管培生offer #

34747次浏览 415人参与

# 如果有时光机，你最想去到哪个年纪？ #