[!WARNING] pigauto is experimental — use at your own risk. It needs further validation. Point estimates are the supported claim. Prediction intervals are nominal held-out diagnostics, not package-wide certification; covariance routes have focused-test evidence only.
pigauto fills missing species trait values using a phylogenetic tree, other traits, and optional covariates. It supports continuous, count, binary, categorical, ordinal, proportion, zero-inflated-count, and compositional traits.
Installation
install.packages("pigauto")
torch::install_torch() # once, after installing the packageStart here
Put a CSV trait table (with a unique species column) and a Newick or NEXUS tree beside your R script. Then use this six-line journey:
library(pigauto)
traits <- read_traits("traits.csv")
tree <- read_tree("tree.nwk")
check_pigauto(traits, tree)
result <- impute(traits, tree)
completed <- completed_data(result)
pigauto_report(result)check_pigauto() runs before fitting: it reports input errors, species/tree matching, trait declarations, runtime availability, and size. Resolve an "error" status before calling impute(); a fully observed target table does not need fitting, so use cross_validate() instead.
The output roles are deliberately separate:
-
completedis the input-shaped trait data with observed values retained and modeled missing cells filled. -
result$predictioncontains all-cell diagnostic predictions, not a second completed dataset. -
result$prediction$seis type-dependent uncertainty; discrete values are not Gaussian standard errors. Conformal bounds are nominal held-out diagnostics. - A
pigauto_resultdoes not itself authorize downstream inference. The supported pooling entries begin withmulti_impute_analysis()in its documented narrow regime or, for continuous traits,multi_impute(draws_method = "posterior").multi_impute()’s default,draws_method = "auto", picks"posterior"whenever the data allow it and otherwise says why it fell back to non-poolable conformal draws.
Tiny installed example inputs are available at system.file("extdata", "novice_traits.csv", package = "pigauto") and system.file("extdata", "novice_tree.nwk", package = "pigauto"). The matching installed six-expression script is system.file("examples", "novice-workflow.R", package = "pigauto").
Copy-ready data recipes
Integer-valued continuous measurements
An integer column is otherwise read as a count. Declare a measurement stored as whole numbers explicitly.
Proportions and zero-inflated counts
Declare proportions only when their numeric 0–1 scale is semantically a proportion. Declare zi_count only when zero is a separate process from the positive count magnitude.
Compositions
Each declared composition is one trait: every row must have all components or all components missing, and observed rows must sum to one.
Repeated observations
read_traits() requires unique species names. For repeated observations, read the CSV directly and supply species_col.
traits_obs <- read.csv("traits_repeated.csv", stringsAsFactors = FALSE)
check_pigauto(traits_obs, tree, species_col = "species")
result <- impute(traits_obs, tree, species_col = "species")Reconciling species names
Inspect the structured species report before changing data:
check <- check_pigauto(traits, tree)
check$species$matched
check$species$data_only
check$species$tree_onlydata_only rows remain in completed but are not modeled or filled. tree_only tips are internal all-missing rows used only while fitting.
Defaults, and when to change them
impute(traits, tree) with its defaults is the starting point. The table says which route to take for each need; reach for another route only for the need in the left-hand column.
| You want | Use | Default? |
|---|---|---|
| Filled-in traits | impute(traits, tree) |
Yes |
| Uncertainty for one imputed value | the per-cell conformal interval in result$prediction
|
Yes |
| A downstream model (regression, PGLS, mixed model) that accounts for imputation |
multi_impute(traits, tree), then with_imputations() and pool_mi()
|
Yes, when every trait is continuous, with one row per species and no covariates: the default draws_method = "auto" then uses posterior draws |
One incomplete covariate in lm, glm or lmer
|
multi_impute_analysis() |
No |
For other data, draws_method = "auto" falls back to "conformal" draws and says so. Those draws, like "mc_dropout", describe the spread of plausible values but are not proper multiple imputations: in a simulation of a phylogenetic regression they biased the pooled slope towards zero, and with_imputations() refuses them. Bayesian downstream fits (brms, MCMCglmm) are combined by stacking posterior draws, for example with brms::brm_multiple(), not with pool_mi(). Worked examples for nlme, lme4, glmmTMB, drmTMB, gllvmTMB and brms are in the multiple imputation article.
Advanced controls
The default journey above is the supported starting point. These controls are for an explicit technical workflow and are kept separate from it.
# Solver, EM, and refinement controls
result <- impute(
traits, tree,
joint_solver = "rphylopars",
em_iterations = 2L,
joint_refine_iter = 1L,
conformal_split_val = TRUE
)
# Prediction route: the default "auto" chooses per trait between the
# cross-trait "exact" conditional and "per_column"; force one with
result <- impute(traits, tree, predict_method = "exact")See ?impute, ?fit_baseline, and the articles for all calibration, refinement, and prediction controls.
