Python 熊猫分组并在列表中获取dict_Python_Python 3.x_Pandas_Dictionary_Pandas Groupby

Python 熊猫分组并在列表中获取dict

python python-3.x pandas dictionary

Python 熊猫分组并在列表中获取dict,python,python-3.x,pandas,dictionary,pandas-groupby,Python,Python 3.x,Pandas,Dictionary,Pandas Groupby,我试图提取分组行数据，以使用值将其与另一个文件中的标签颜色一起打印我的数据框如下所示 df = pd.DataFrame({'x': [1, 4, 5], 'y': [3, 2, 5], 'label': [1.0, 1.0, 2.0]}) x y label 0 1 3 1.0 1 4 2 1.0 2 5 5 2.0 我想得到一组标签列表，如 {'1.0': [{'index': 0, 'x': 1, 'y': 3}, {'index'

我试图提取分组行数据，以使用值将其与另一个文件中的标签颜色一起打印

我的数据框如下所示

df = pd.DataFrame({'x': [1, 4, 5], 'y': [3, 2, 5], 'label': [1.0, 1.0, 2.0]})

    x   y   label
0   1   3   1.0
1   4   2   1.0
2   5   5   2.0

我想得到一组标签列表，如

{'1.0': [{'index': 0, 'x': 1, 'y': 3}, {'index': 1, 'x': 4, 'y': 2}],
 '2.0': [{'index': 2, 'x': 5, 'y': 5}]}

如何做到这一点？

您可以使用和：

df = pd.DataFrame({'x': [1, 4, 5], 'y': [3, 2, 5], 'label': [1.0, 1.0, 2.0]})
df['index'] = df.index
df
   label  x  y  index
0    1.0  1  3      0
1    1.0  4  2      1
2    2.0  5  5      2

df['dict']=df[['x','y','index']].to_dict("records")
df
   label  x  y  index                             dict
0    1.0  1  3      0  {u'y': 3, u'x': 1, u'index': 0}
1    1.0  4  2      1  {u'y': 2, u'x': 4, u'index': 1}
2    2.0  5  5      2  {u'y': 5, u'x': 5, u'index': 2}

df = df[['label','dict']]
df['label'] = df['label'].apply(str) #Converting integer column 'label' to string
df = df.groupby('label')['dict'].apply(list) 
desired_dict = df.to_dict()
desired_dict 
    {'1.0': [{'index': 0, 'x': 1, 'y': 3}, {'index': 1, 'x': 4, 'y': 2}],
     '2.0': [{'index': 2, 'x': 5, 'y': 5}]}

itertuples返回命名元组以在数据帧上迭代：

for row in df.itertuples():
    print(row)
Pandas(Index=0, x=1, y=3, label=1.0)
Pandas(Index=1, x=4, y=2, label=1.0)
Pandas(Index=2, x=5, y=5, label=2.0)

因此，利用这一点：

from collections import defaultdict
dictionary = defaultdict(list)
for row in df.itertuples():
    dummy['x'] = row.x
    dummy['y'] = row.y
    dummy['index'] = row.Index
    dictionary[row.label].append(dummy)

dict(dictionary)
> {1.0: [{'x': 1, 'y': 3, 'index': 0}, {'x': 4, 'y': 2, 'index': 1}],
 2.0: [{'x': 5, 'y': 5, 'index': 2}]}

最快的解决方案几乎就是@cph_sto提供的解决方案

>>> df.reset_index().to_dict('records')
[{'index': 0.0, 'label': 1.0, 'x': 1.0, 'y': 3.0}, {'index': 1.0, 'label': 1.0, 'x': 4.0, 'y': 2.0}, {'index': 2.0, 'label': 2.0, 'x': 5.0, 'y': 5.0}]

也就是说，将索引转换为常规列，然后将

记录的版本应用于dict
。另一个感兴趣的选择：
>>> df.to_dict('index')
{0: {'label': 1.0, 'x': 1.0, 'y': 3.0}, 1: {'label': 1.0, 'x': 4.0, 'y': 2.0}, 2: {'label': 2.0, 'x': 5.0, 'y': 5.0}}

有关更多信息，请查看帮助。
您可以使用：
结果:
print(dd)

defaultdict(list,
            {1.0: [{'index': 0.0, 'x': 1.0, 'y': 3.0, 'label': 1.0},
                   {'index': 1.0, 'x': 4.0, 'y': 2.0, 'label': 1.0}],
             2.0: [{'index': 2.0, 'x': 5.0, 'y': 5.0, 'label': 2.0}]})

一般来说，没有必要转换回常规的dict
，因为defaultdict
是dict的一个子类，感谢它的工作。我知道defaultdict和dict一样可用。谢谢。据我所知，我会检查熊猫的口述文件。谢谢。你的答案正是我想要的。我使用itertuple@rootpetit很高兴我能帮上忙：）请接受回答您问题的答案：）感谢教学过程。了解熊猫对我帮助很大。
print(dd)

defaultdict(list,
            {1.0: [{'index': 0.0, 'x': 1.0, 'y': 3.0, 'label': 1.0},
                   {'index': 1.0, 'x': 4.0, 'y': 2.0, 'label': 1.0}],
             2.0: [{'index': 2.0, 'x': 5.0, 'y': 5.0, 'label': 2.0}]})