我有一個這種結構的大資料框:
structure(list(Date = structure(c(18220, 18220, 18220,
18955, 19110, 19110, 18514, 18514, 18892, 18647, 18647, 18528,
18822, 18822, 18822, 18822, 18745, 18745, 18745), class = "Date"),
Type = c("Main", "Main", "Main",
"Main", "Secondary", "Tri-Annual", "Main",
"Main", "Tri-Annual", "Main", "Syndication",
"Main", "Tri-Annual", "Syndication", "Secondary",
"Main", "Main", "Tri-Annual", "Secondary"
), Value = c(4000, 2250, 2250, 1100, 1800, 12000,
8000, 9000, 10000, 6500, 7000, 6500, 7000, 4250, 5500, 2500,
6000, 6000, 4500), buckets = c("Long", "Long", "Long",
"Long", "Long", "Medium", "Medium", "Long", "Medium", "Long",
"Long", "Long", "Long", "Long", "Long", "Long", "Long", "Long",
"Long")), row.names = c(NA, -19L), class = c("tbl_df", "tbl",
"data.frame"))
我正在嘗試使用以下方法按型別和日期聚合資料:
df <- df %>% group_by(Date, Type) %>% dplyr::summarise(Aggregated = sum(df$Value))
但是,它不是每天給出每個型別的值的總和,而是將所有值加在一起(無論是哪一天)并將其發布在聚合列中。有誰知道這有什么問題?
uj5u.com熱心網友回復:
檢查變數呼叫的使用。
如果您使用 df$Value 作為summarise()呼叫的一部分,則總和將建立在該物件的向量上。這意味著您“離開”管道并尋找一個物件(恰好是相同的,即 df)并獲取其列值),而不是匯總分組的 tibble/dataframe。
df %>%
group_by(Date, Type) %>%
summarise(N = n() # just to count we have several entries per Date and Type
, aggregated_old = sum(df$Value) # you call the dataframe column
, aggregated_new = sum(Value) # do not use object df, just column name
)
# A tibble: 16 × 5
# Groups: Date [9]
Date Type N aggregated_old aggregated_new
<date> <chr> <int> <dbl> <dbl>
1 2019-11-20 Main 3 106150 8500
2 2020-09-09 Main 2 106150 17000
3 2020-09-23 Main 1 106150 6500
4 2021-01-20 Main 1 106150 6500
5 2021-01-20 Syndication 1 106150 7000
6 2021-04-28 Main 1 106150 6000
7 2021-04-28 Secondary 1 106150 4500
8 2021-04-28 Tri-Annual 1 106150 6000
9 2021-07-14 Main 1 106150 2500
10 2021-07-14 Secondary 1 106150 5500
11 2021-07-14 Syndication 1 106150 4250
12 2021-07-14 Tri-Annual 1 106150 7000
13 2021-09-22 Tri-Annual 1 106150 10000
14 2021-11-24 Main 1 106150 1100
15 2022-04-28 Secondary 1 106150 1800
16 2022-04-28 Tri-Annual 1 106150 12000
轉載請註明出處,本文鏈接:https://www.uj5u.com/houduan/524757.html
標籤:r数据框dplyr
