Need to arrange dataframe as a summary dataframe using R

Question

I am using the following dataframe in R.

dput:

structure(list(uid = c("K-1", "K-1", 
"K-2", "K-3", "K-4", "K-5", 
"K-6", "K-7", "K-8", "K-9", 
"K-10", "K-11", "K-12", "K-13", 
"K-14"), Date = c("2020-03-16 12:11:33", "2020-03-16 12:11:33", 
"2020-03-16 06:13:55", "2020-03-16 10:03:43", "2020-03-16 12:37:09", 
"2020-03-16 06:40:24", "2020-03-16 09:46:45", "2020-03-16 12:07:44", 
"2020-03-16 14:09:51", "2020-03-16 09:19:23", "2020-03-16 09:07:37", 
"2020-03-16 11:48:34", "2020-03-16 06:23:24", "2020-03-16 04:39:03", 
"2020-03-16 04:59:13"), batch_no = c(7, 7, 8, 9, 9, 8, 
7, 6, 7, 9, 8, 8, 7, 7, 7), marking = c("S1", "S1", "S2", 
"SE_hold1", "SD_hold1", "SD_hold2", "S3", "S3", "", "SA_hold3", "S1", "S1", "S2", 
"S3", "S3"), seq = c("FRD", 
"FHL", NA, NA, NA, NA, NA, NA, "ABC", NA, NA, NA, NA, "DEF", NA)), .Names = c("uid", 
"Date", "batch_no", "marking",
"seq"), row.names = c(NA, 15L), class = "data.frame")






uid     Date                  batch_no       marking       seq
K-1     16/03/2020  12:11:33  7              S1            FRD
K-1     16/03/2020  12:11:33  7              S1            FHL
K-2     16/03/2020  12:11:33  8              SE_hold1      ABC
K-3     16/03/2020  12:11:33  9              SD_hold2      DEF
K-4     16/03/2020  12:11:33  8              S1            XYZ
K-5     16/03/2020  12:11:33                 NA            ABC
K-6     16/03/2020  12:11:33  7                            ZZZ
K-7     16/03/2020  12:11:33  NA             S2            NA
K-8     16/03/2020  12:11:33  6              S3            FRD

The seq column will have eight unique value including NA, not necessary the all 8 values are available for every day's date.
batch_no will have six unique values including NA and blank, not necessary the all six values are available for every day's date.
The marking column will have ~ 25 unique value but need to consider values with suffix _hold# as Hold after that there would be six unique value including blank and NA.

The requirement is to merge the dcast dataframe in the following order to have a single view summary for an analysis.

I want to keep all the unique values static in the code, so that if the particular value is not available for a particular date I'll get 0 or - in summary table.

Desired Output:

seq      count  percentage   Marking     count     Percentage     batch_no   count    Percentage
FRD      1      12.50%       S1          2         25.00%         6          1        12.50%
FHL      1      12.50%       S2          1         12.50%         7          2        25.00%
ABC      2      25.00%       S3          1         12.50%         8          2        25.00%
DEF      1      12.50%       Hold        2         25.00%         9          1        12.50%
XYZ      1      12.50%       NA          1         12.50%         NA         1        12.50%
ZZZ      1      12.50%       (Blank)     1         12.50%         (Blank)    1        12.50%
FRD      1      12.50%         -         -           -             -         -           -
NA       1      12.50%         -         -           -             -         -           -
(Blank)  0      0.00%          -         -           -             -         -           -
Total    8      112.50%        -         8         100.00%         -         8         100.00%

For seq we have % > 100 because of double counting of same uid for value FRD and FHL. That is the accepted scenario. In Total will have only distinct count of uid.

I'm using below-mentioned code by SO, but couldn't get the desired output.

df = df_original %>% 
  mutate(marking = if_else(str_detect(marking,"hold"),"Hold", marking)) %>% 
  mutate_at(vars(c("seq", "batch_no", "marking")), forcats::fct_explicit_na, na_level = "(Blank)") 

## You Need to do something similar with vectors of the possible values
df_combinations = purrr::cross_df(list(seq = df$seq %>% unique(),
                     batch_no = df$batch_no %>% unique(),
                     marking = df$marking %>% unique()))

df_all_combination = df_combinations %>% 
  left_join(df, by = c("seq", "batch_no", "marking")) %>% 
  group_by(seq, batch_no, marking) %>% 
  summarise(count = n())

Could you make your problem reproducible by sharing a sample of your data so others can help (please do not use `str()`, `head()` or screenshot)? You can use the [`reprex`](https://reprex.tidyverse.org/articles/articles/magic-reprex.html) and [`datapasta`](https://cran.r-project.org/web/packages/datapasta/vignettes/how-to-datapasta.html) packages to assist you with that. See also [Help me Help you](https://speakerdeck.com/jennybc/reprex-help-me-help-you?slide=5) & [How to make a great R reproducible example?](https://stackoverflow.com/q/5963269) — Tung, Apr 10 '20 at 11:50

hello_friend · Accepted Answer · 2020-04-10T15:12:41.590

2

Base R solution (note: I'm not entirely sure I understand your question):

  # Function to summarise each of the vectors required: summariser => function
summariser <- function(vec) {
  within(unique(data.frame(
    vec = vec,
    counter = as.numeric(ifelse(is.na(vec), sum(is.na(vec)),
                     ave(vec, vec, FUN = length))), stringsAsFactors = FALSE
  )),
  {
    perc = paste0(round(counter / sum(counter) * 100, 2), "%")
  })
}
# Vectors to summarise: vecs_to_summarise => character vector
vecs_to_summarise <- c("seq", "marking", "batch_no")

# Create an empty list in order to allocate some memory: df_list => list
df_list <- vector("list", length(vecs_to_summarise))

# Apply the summariser function to each of the vectors required: df_list => list of dfs
df_list <- lapply(df[,vecs_to_summarise], summariser)

# Rename the vectors of each data.frame in the list: df_list => list of dfs: 
df_list <- lapply(seq_along(df_list), function(i) {
  names(df_list[[i]]) <- gsub("_vec", "",
                              paste(names(df_list[i]), names(df_list[[i]]), sep = "_"))
  return(df_list[[i]])
})


# Determine the number of rows of the maximum data.frame: numeric scalar 
max_df_length <- max(sapply(df_list, nrow))

# Extend each data.frame to be the same length (pad with NAs if necessary): df_list => list
df_list <- lapply(seq_along(df_list), function(i){
  y <- data.frame(df_list[[i]][rep(seq_len(nrow(df_list[[1]])), each = 1),])
  y[1:(nrow(y)),] <- NA
  y <- y[1:(max_df_length - nrow(df_list[[i]])),]
  if(length(y) > 0){
    x <- data.frame(rbind(df_list[[i]], y)[1:max_df_length,])
  }else{
      x <- data.frame(df_list[[i]][1:max_df_length,])
      }
  return(x)
  }
)

# Bind the data.frames in the list into a single df: analysed_df => data.frame
analysed_df <- do.call("cbind", df_list)

Data:

df <- structure(list(uid = c("K-1", "K-1", "K-2", "K-3", "K-4", "K-5", 
    "K-6", "K-7", "K-8"), Date = structure(c(1584321093, 1584321093, 
    1584321093, 1584321093, 1584321093, 1584321093, 1584321093, 1584321093, 
    1584321093), class = c("POSIXct", "POSIXt"), tzone = ""), batch_no = c(7L, 
    7L, 8L, 9L, 8L, NA, 7L, NA, 6L), marking = c("S1", "S1", "SE_hold1", 
    "SD_hold2", "S1", NA, NA, "S2", "S3"), seq = c("FRD", "FHL", 
    "ABC", "DEF", "XYZ", "ABC", "ZZZ", NA, "FRD")), row.names = c(NA, 
    -9L), class = "data.frame")

edited Apr 10 '20 at 15:12

answered Apr 10 '20 at 12:35

hello_friend

5,682
1
11
15

`ordered_df` is not unique, I need the count basis unique value only :( – Sophia Wilson Apr 10 '20 at 12:47
I'm updating your answer with the output that I get. I'm total number of rows in the desired output. – Sophia Wilson Apr 10 '20 at 13:00
In fact with sample data I'm getting duplicate in `batch_no` and `marking`. – Sophia Wilson Apr 10 '20 at 13:15
@SophiaWilson very close give me a few mins – hello_friend Apr 10 '20 at 14:42
@SophiaWilson done please see edited answer. Also upvote and accept it if it does what you want. Also please learn how to post questions on here. See here: https://www.r-bloggers.com/three-tips-for-posting-good-questions-to-r-help-and-stack-overflow/ – hello_friend Apr 10 '20 at 15:01
Thank, it worked. Only `Total` column is missing that I'll add. – Sophia Wilson Apr 10 '20 at 15:09
@SophiaWilson just edited to rename cols_to_summarise => vecs_to_summarise – hello_friend Apr 10 '20 at 15:13
One small help, how can I keep the variables of `seq`, `marking` and `batch_no` static so that if for a particular date if a particular variable is missing then I can show it as 0. – Sophia Wilson Apr 10 '20 at 16:05
@SophiaWilson please edit example data to reflect your request. – hello_friend Apr 11 '20 at 03:17
Hi, I just want to know how can i kept some variable static if that are not available for any particular date. For Example, if the `S2`, `S3` (from `marking` or `8`,`9` (from `batch_no`) are not available for a date x then I want the palce holder for these variables in `analysed_df` with value = 0. – Sophia Wilson Apr 11 '20 at 11:04
How to add total column at last row, if possible for both count & percentage. – Sophia Wilson Apr 17 '20 at 05:51
@SophiaWilson Have you posted the other question ? – hello_friend Apr 17 '20 at 16:38
Please help me add the last low as total for both `count` and `percentage`. – Sophia Wilson Apr 18 '20 at 05:40
@SophiaWilson I answered your bountied question with the solution to that problem. – hello_friend Apr 18 '20 at 05:40
@SophiaWilson please upvote and accept my solution here: https://stackoverflow.com/questions/61166860/how-to-keep-some-variables-static-for-the-summary-in-r/61284441#61284441 – hello_friend Apr 18 '20 at 05:43
Done, This code was comparatively easy to understand that's why I asked. – Sophia Wilson Apr 18 '20 at 05:45
@SophiaWilson thats because what it was doing was more R like, still keeping (mostly) appropriated type vectors, with tidy(-ish) data.frames. Also it was doing a lot less than you what your desired result asked of it. The new solution meets the spec of your reporting criteria 100%. – hello_friend Apr 18 '20 at 05:48

Need to arrange dataframe as a summary dataframe using R

1 Answers1