<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Papers on statistical.systems</title>
    <link>https://statistical.systems/papers/</link>
    <description>Recent content in Papers on statistical.systems</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en-us</language>
    <lastBuildDate>Thu, 20 Aug 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://statistical.systems/papers/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Methylation Changes in the Peripheral Blood of Filipinos with Type 2 Diabetes Suggest Spurious Transcription Initiation at TXNIP</title>
      <link></link>
      <pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate>
      
      <guid></guid>
      <description>&lt;p&gt;While much work has been done in associating differentially methylated positions (DMPs) to type 2 diabetes (T2D) across different populations, not much attention has been placed on identifying its possible functional consequences. We explored methylation changes in the peripheral blood of Filipinos with T2D and identified 177 associated DMPs. Most of these DMPs were associated with genes involved in metabolism, inflammation and the cell cycle. Three of these DMPs map to the TXNIP gene body, replicating previous findings from epigenome-wide association studies (EWAS) of T2D. The TXNIP downmethylation coincided with increased transcription at the 3′ UTR, H3K36me3 histone markings and Sp1 binding, suggesting spurious transcription initiation at the TXNIP 3′ UTR as a functional consequence of T2D methylation changes. We also explored potential epigenetic determinants to increased incidence of T2D in Filipino immigrants in the USA and found three DMPs associated with the interaction of T2D and immigration. Two of these DMPs were located near MAP2K7 and PRMT1, which may point towards dysregulated stress response and inflammation as a contributing factor to T2D among Filipino immigrants.&lt;/p&gt;</description>
      <content:encoded><![CDATA[<p>While much work has been done in associating differentially methylated positions (DMPs) to type 2 diabetes (T2D) across different populations, not much attention has been placed on identifying its possible functional consequences. We explored methylation changes in the peripheral blood of Filipinos with T2D and identified 177 associated DMPs. Most of these DMPs were associated with genes involved in metabolism, inflammation and the cell cycle. Three of these DMPs map to the TXNIP gene body, replicating previous findings from epigenome-wide association studies (EWAS) of T2D. The TXNIP downmethylation coincided with increased transcription at the 3′ UTR, H3K36me3 histone markings and Sp1 binding, suggesting spurious transcription initiation at the TXNIP 3′ UTR as a functional consequence of T2D methylation changes. We also explored potential epigenetic determinants to increased incidence of T2D in Filipino immigrants in the USA and found three DMPs associated with the interaction of T2D and immigration. Two of these DMPs were located near MAP2K7 and PRMT1, which may point towards dysregulated stress response and inflammation as a contributing factor to T2D among Filipino immigrants.</p>
]]></content:encoded>
    </item>
    
    <item>
      <title>Correlation-Adjusted Simultaneous Testing for Ultra High-Dimensional Grouped Data</title>
      <link></link>
      <pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate>
      
      <guid></guid>
      <description>&lt;p&gt;Epigenetics plays a crucial role in understanding the underlying molecular processes of several types of cancer as well as the determination of innovative therapeutic tools. To investigate the complex interplay between genetics and environment, we develop a novel procedure to identify differentially methylated probes (DMPs) among cases and controls. Statistically, this translates to an ultra high-dimensional testing problem with sparse signals and an inherent grouping structure. When the total number of variables being tested is massive and typically exhibits some degree of dependence, existing group-wise multiple comparisons adjustment methods lead to inflated false discoveries. We propose a class of Correlation-Adjusted Simultaneous Testing (CAST) procedures incorporating the general dependence among probes within and between genes to control the false discovery rate (FDR). Simulations demonstrate that CASTs have superior empirical power while maintaining the FDR compared to the benchmark group-wise. Moreover, while the benchmark fails to control FDR for small-sized grouped correlated data, CAST exhibits robustness in controlling FDR across varying group sizes. In bladder cancer data, the proposed CAST method confirms some existing differentially methylated probes implicated with the disease (Langevin, et. al., 2014). However, CAST was able to detect novel DMPs that the previous study (Langevin, et. al., 2014) failed to identify. The CAST method can accurately identify significant potential biomarkers and facilitates informed decision-making aligned with precision medicine in the context of complex data analysis.&lt;/p&gt;</description>
      <content:encoded><![CDATA[<p>Epigenetics plays a crucial role in understanding the underlying molecular processes of several types of cancer as well as the determination of innovative therapeutic tools. To investigate the complex interplay between genetics and environment, we develop a novel procedure to identify differentially methylated probes (DMPs) among cases and controls. Statistically, this translates to an ultra high-dimensional testing problem with sparse signals and an inherent grouping structure. When the total number of variables being tested is massive and typically exhibits some degree of dependence, existing group-wise multiple comparisons adjustment methods lead to inflated false discoveries. We propose a class of Correlation-Adjusted Simultaneous Testing (CAST) procedures incorporating the general dependence among probes within and between genes to control the false discovery rate (FDR). Simulations demonstrate that CASTs have superior empirical power while maintaining the FDR compared to the benchmark group-wise. Moreover, while the benchmark fails to control FDR for small-sized grouped correlated data, CAST exhibits robustness in controlling FDR across varying group sizes. In bladder cancer data, the proposed CAST method confirms some existing differentially methylated probes implicated with the disease (Langevin, et. al., 2014). However, CAST was able to detect novel DMPs that the previous study (Langevin, et. al., 2014) failed to identify. The CAST method can accurately identify significant potential biomarkers and facilitates informed decision-making aligned with precision medicine in the context of complex data analysis.</p>
]]></content:encoded>
    </item>
    
    <item>
      <title>Empirical Null Estimation using Zero-inflated Discrete Mixture Distributions and its Application to Protein Domain Data</title>
      <link></link>
      <pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate>
      
      <guid></guid>
      <description>&lt;p&gt;In recent mutation studies, analyses based on protein domain positions are gaining popularity over gene-centric approaches since the latter have limitations in considering the functional context that the position of the mutation provides. This presents a large-scale simultaneous inference problem, with hundreds of hypothesis tests to consider at the same time. This article aims to select significant mutation counts while controlling a given level of Type I error via False Discovery Rate (FDR) procedures. One main assumption is that the mutation counts follow a zero-inflated model in order to account for the true zeros in the count model and the excess zeros. The class of models considered is the Zero-inflated Generalized Poisson (ZIGP) distribution. Furthermore, we assumed that there exists a cut-off value such that smaller counts than this value are generated from the null distribution. We present several data-dependent methods to determine the cut-off value. We also consider a two-stage procedure based on screening process so that the number of mutations exceeding a certain value should be considered as significant mutations. Simulated and protein domain data sets are used to illustrate this procedure in estimation of the empirical null using a mixture of discrete distributions. Overall, while maintaining control of the FDR, the proposed two-stage testing procedure has superior empirical power.&lt;/p&gt;</description>
      <content:encoded><![CDATA[<p>In recent mutation studies, analyses based on protein domain positions are gaining popularity over gene-centric approaches since the latter have limitations in considering the functional context that the position of the mutation provides. This presents a large-scale simultaneous inference problem, with hundreds of hypothesis tests to consider at the same time. This article aims to select significant mutation counts while controlling a given level of Type I error via False Discovery Rate (FDR) procedures. One main assumption is that the mutation counts follow a zero-inflated model in order to account for the true zeros in the count model and the excess zeros. The class of models considered is the Zero-inflated Generalized Poisson (ZIGP) distribution. Furthermore, we assumed that there exists a cut-off value such that smaller counts than this value are generated from the null distribution. We present several data-dependent methods to determine the cut-off value. We also consider a two-stage procedure based on screening process so that the number of mutations exceeding a certain value should be considered as significant mutations. Simulated and protein domain data sets are used to illustrate this procedure in estimation of the empirical null using a mixture of discrete distributions. Overall, while maintaining control of the FDR, the proposed two-stage testing procedure has superior empirical power.</p>
]]></content:encoded>
    </item>
    
    <item>
      <title>Bayesian Local False Discovery Rate for Sparse Count Data with Application to the Discovery of Hotspots in Protein Domains</title>
      <link></link>
      <pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate>
      
      <guid></guid>
      <description>&lt;p&gt;In cancer research at the molecular level, it is critical to understand which somatic mutations play an important role in the initiation or progression of cancer. Recently, studying cancer somatic variants at the protein domain level is an important area for uncovering functionally related somatic mutations. The main issue is to find the protein domain hotspots which have significantly high frequency of mutations. Multiple testing procedures are commonly used to identify hotspots; however, when data is not large enough, existing methods produce unreliable results with failure in controlling a given type I error rate. We propose multiple testing procedures, based on Bayesian local false discovery rate, for sparse count data and apply it in the identification of clusters of somatic mutations across entire gene families using protein domain models. In multiple testing for count data, it is not clear what kind of the null distribution should be admitted. In our proposed algorithms, we implement the zero assumption in the context of Bayesian methods to identify the null distribution for count data rather than using any theoretical null distribution. Furthermore, we also address different types of modeling of alternative distributions. The proposed fully Bayesian models are efficient when the number of count data is small (50 ≤ N &amp;lt; 200) while the local false discovery rate procedures, based on the empirical Bayes, is desirable for a large number of data (N &amp;gt; 800). We provide numerical studies to show that the proposed fully Bayesian methods can control a given level of false discovery rate for small number of positions while existing approaches based on nonparametric empirical Bayes fail in controlling a false discovery rate. In addition, we present real data examples of protein domain data to select hotspots in protein domain data.&lt;/p&gt;</description>
      <content:encoded><![CDATA[<p>In cancer research at the molecular level, it is critical to understand which somatic mutations play an important role in the initiation or progression of cancer. Recently, studying cancer somatic variants at the protein domain level is an important area for uncovering functionally related somatic mutations. The main issue is to find the protein domain hotspots which have significantly high frequency of mutations. Multiple testing procedures are commonly used to identify hotspots; however, when data is not large enough, existing methods produce unreliable results with failure in controlling a given type I error rate. We propose multiple testing procedures, based on Bayesian local false discovery rate, for sparse count data and apply it in the identification of clusters of somatic mutations across entire gene families using protein domain models. In multiple testing for count data, it is not clear what kind of the null distribution should be admitted. In our proposed algorithms, we implement the zero assumption in the context of Bayesian methods to identify the null distribution for count data rather than using any theoretical null distribution. Furthermore, we also address different types of modeling of alternative distributions. The proposed fully Bayesian models are efficient when the number of count data is small (50 ≤ N &lt; 200) while the local false discovery rate procedures, based on the empirical Bayes, is desirable for a large number of data (N &gt; 800). We provide numerical studies to show that the proposed fully Bayesian methods can control a given level of false discovery rate for small number of positions while existing approaches based on nonparametric empirical Bayes fail in controlling a false discovery rate. In addition, we present real data examples of protein domain data to select hotspots in protein domain data.</p>
]]></content:encoded>
    </item>
    
    <item>
      <title>Oncodomains: A Protein Domain-Centric Framework for Analyzing Rare Variants in Tumor Samples</title>
      <link></link>
      <pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate>
      
      <guid></guid>
      <description>&lt;p&gt;The fight against cancer is hindered by its highly heterogeneous nature. Genome-wide sequencing studies have shown that individual malignancies contain many mutations that range from those commonly found in tumor genomes to rare somatic variants present only in a small fraction of lesions. Such rare somatic variants dominate the landscape of genomic mutations in cancer, yet efforts to correlate somatic mutations found in one or few individuals with functional roles have been largely unsuccessful. Traditional methods for identifying somatic variants that drive cancer are &amp;ldquo;gene-centric&amp;rdquo; in that they consider only somatic variants within a particular gene and make no comparison to other similar genes in the same family that may play a similar role in cancer. In this work, we present oncodomain hotspots, a new &amp;ldquo;domain-centric&amp;rdquo; method for identifying clusters of somatic mutations across entire gene families using protein domain models. Our analysis confirms that our approach creates a framework for leveraging structural and functional information encapsulated by protein domains into the analysis of somatic variants in cancer, enabling the assessment of even rare somatic variants by comparison to similar genes. Our results reveal a vast landscape of somatic variants that act at the level of domain families altering pathways known to be involved with cancer such as protein phosphorylation, signaling, gene regulation, and cell metabolism. Due to oncodomain hotspots&amp;rsquo; unique ability to assess rare variants, we expect our method to become an important tool for the analysis of sequenced tumor genomes, complementing existing methods.&lt;/p&gt;</description>
      <content:encoded><![CDATA[<p>The fight against cancer is hindered by its highly heterogeneous nature. Genome-wide sequencing studies have shown that individual malignancies contain many mutations that range from those commonly found in tumor genomes to rare somatic variants present only in a small fraction of lesions. Such rare somatic variants dominate the landscape of genomic mutations in cancer, yet efforts to correlate somatic mutations found in one or few individuals with functional roles have been largely unsuccessful. Traditional methods for identifying somatic variants that drive cancer are &ldquo;gene-centric&rdquo; in that they consider only somatic variants within a particular gene and make no comparison to other similar genes in the same family that may play a similar role in cancer. In this work, we present oncodomain hotspots, a new &ldquo;domain-centric&rdquo; method for identifying clusters of somatic mutations across entire gene families using protein domain models. Our analysis confirms that our approach creates a framework for leveraging structural and functional information encapsulated by protein domains into the analysis of somatic variants in cancer, enabling the assessment of even rare somatic variants by comparison to similar genes. Our results reveal a vast landscape of somatic variants that act at the level of domain families altering pathways known to be involved with cancer such as protein phosphorylation, signaling, gene regulation, and cell metabolism. Due to oncodomain hotspots&rsquo; unique ability to assess rare variants, we expect our method to become an important tool for the analysis of sequenced tumor genomes, complementing existing methods.</p>
]]></content:encoded>
    </item>
    
    <item>
      <title>Bayesian Variable Selection using Knockoffs with Applications to Genomics</title>
      <link></link>
      <pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate>
      
      <guid></guid>
      <description>&lt;p&gt;Given the costliness of HIV drug therapy research, it is important not only to maximize true positive rate (TPR) by identifying which genetic markers are related to drug resistance, but also to minimize false discovery rate (FDR) by reducing the number of incorrect markers unrelated to drug resistance. In this study, we propose a multiple testing procedure that unifies key concepts in computational statistics, namely Model-free Knockoffs, Bayesian variable selection, and the local false discovery rate. We develop an algorithm that utilizes the augmented data-Knockoff matrix and implement Bayesian Lasso. We then identify signals using test statistics based on Markov Chain Monte Carlo outputs and local false discovery rate. We test our proposed methods against non-bayesian methods such as Benjamini–Hochberg (BHq) and Lasso regression in terms TPR and FDR. Using numerical studies, we show the proposed method yields lower FDR compared to BHq and Lasso for certain cases, such as for low and equi-dimensional cases. We also discuss an application to an HIV-1 data set, which aims to be applied analyzing genetic markers linked to drug resistant HIV in the Philippines in future work.&lt;/p&gt;</description>
      <content:encoded><![CDATA[<p>Given the costliness of HIV drug therapy research, it is important not only to maximize true positive rate (TPR) by identifying which genetic markers are related to drug resistance, but also to minimize false discovery rate (FDR) by reducing the number of incorrect markers unrelated to drug resistance. In this study, we propose a multiple testing procedure that unifies key concepts in computational statistics, namely Model-free Knockoffs, Bayesian variable selection, and the local false discovery rate. We develop an algorithm that utilizes the augmented data-Knockoff matrix and implement Bayesian Lasso. We then identify signals using test statistics based on Markov Chain Monte Carlo outputs and local false discovery rate. We test our proposed methods against non-bayesian methods such as Benjamini–Hochberg (BHq) and Lasso regression in terms TPR and FDR. Using numerical studies, we show the proposed method yields lower FDR compared to BHq and Lasso for certain cases, such as for low and equi-dimensional cases. We also discuss an application to an HIV-1 data set, which aims to be applied analyzing genetic markers linked to drug resistant HIV in the Philippines in future work.</p>
]]></content:encoded>
    </item>
    
    <item>
      <title>Causal Relationships Between Diseases Mined from the Literature Improve the Use of Polygenic Risk Scores</title>
      <link></link>
      <pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate>
      
      <guid></guid>
      <description>&lt;p&gt;&lt;strong&gt;Motivation.&lt;/strong&gt; Identifying causal relations between diseases allows for the study of shared pathways, biological mechanisms, and inter-disease risks. Such causal relations can facilitate the identification of potential disease precursors and candidates for drug re-purposing. However, computational methods often lack access to these causal relations. Few approaches have been developed to automatically extract causal relationships between diseases from unstructured text, but they are often only focused on a small number of diseases, lack validation of the extracted causal relations, or do not make their data available.&lt;/p&gt;</description>
      <content:encoded><![CDATA[<p><strong>Motivation.</strong> Identifying causal relations between diseases allows for the study of shared pathways, biological mechanisms, and inter-disease risks. Such causal relations can facilitate the identification of potential disease precursors and candidates for drug re-purposing. However, computational methods often lack access to these causal relations. Few approaches have been developed to automatically extract causal relationships between diseases from unstructured text, but they are often only focused on a small number of diseases, lack validation of the extracted causal relations, or do not make their data available.</p>
<p><strong>Results.</strong> We automatically mined statements asserting a causal relation between diseases from the scientific literature by leveraging lexical patterns. Following automated mining of causal relations, we mapped the diseases to the International Classification of Diseases (ICD) identifiers to allow the direct application to clinical data. We provide quantitative and qualitative measures to evaluate the mined causal relations and compare to UK Biobank diagnosis data as a completely independent data source. The validated causal associations were used to create a directed acyclic graph that can be used by causal inference frameworks. We demonstrate the utility of our causal network by performing causal inference using the do-calculus, using relations within the graph to construct and improve polygenic risk scores, and disentangle the pleiotropic effects of variants.</p>
<p><strong>Availability and implementation.</strong> The data are available through <a href="https://github.com/bio-ontology-research-group/causal-relations-between-diseases">github.com/bio-ontology-research-group/causal-relations-between-diseases</a>.</p>
]]></content:encoded>
    </item>
    
    <item>
      <title>Ridge Penalization in High-Dimensional Testing with Applications to Imaging Genetics</title>
      <link></link>
      <pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate>
      
      <guid></guid>
      <description>&lt;p&gt;High-dimensionality is ubiquitous in various scientific fields such as imaging genetics, where a deluge of functional and structural data on brain-relevant genetic polymorphisms are investigated. It is crucial to identify which genetic variations are consequential in identifying neurological features of brain connectivity compared to merely random noise. Statistical inference in high-dimensional settings poses multiple challenges involving analytical and computational complexity. A widely implemented strategy in addressing inference goals is penalized inference. In particular, the role of the ridge penalty in high-dimensional prediction and estimation has been actively studied in the past several years. This study focuses on ridge-penalized tests in high-dimensional hypothesis testing problems by proposing and examining a class of methods for choosing the optimal ridge penalty. We present our findings on strategies to improve the statistical power of ridge-penalized tests and what determines the optimal ridge penalty for hypothesis testing. The application of our work to an imaging genetics study and biological research will be presented.&lt;/p&gt;</description>
      <content:encoded><![CDATA[<p>High-dimensionality is ubiquitous in various scientific fields such as imaging genetics, where a deluge of functional and structural data on brain-relevant genetic polymorphisms are investigated. It is crucial to identify which genetic variations are consequential in identifying neurological features of brain connectivity compared to merely random noise. Statistical inference in high-dimensional settings poses multiple challenges involving analytical and computational complexity. A widely implemented strategy in addressing inference goals is penalized inference. In particular, the role of the ridge penalty in high-dimensional prediction and estimation has been actively studied in the past several years. This study focuses on ridge-penalized tests in high-dimensional hypothesis testing problems by proposing and examining a class of methods for choosing the optimal ridge penalty. We present our findings on strategies to improve the statistical power of ridge-penalized tests and what determines the optimal ridge penalty for hypothesis testing. The application of our work to an imaging genetics study and biological research will be presented.</p>
]]></content:encoded>
    </item>
    
    <item>
      <title>Predictive Performance Test based on the Exhaustive Nested Cross-Validation for High-Dimensional Data</title>
      <link></link>
      <pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate>
      
      <guid></guid>
      <description>&lt;p&gt;It is crucial to assess the predictive performance of a model to establish its practicality and relevance in real-world scenarios, particularly for high-dimensional data analysis. Among data splitting or resampling methods, cross-validation (CV) is extensively used for several tasks such as estimating the prediction error, tuning the regularization parameter, and selecting the most suitable predictive model among competing alternatives. The K-fold cross-validation is a popular CV method but its limitation is that the risk estimates are highly dependent on the partitioning of the data (for training and testing). Here, the issues regarding the reproducibility of the K-fold CV estimator are demonstrated in hypothesis testing wherein different partitions lead to notably disparate conclusions. This study presents a novel predictive performance test and valid confidence intervals based on exhaustive nested cross-validation for determining the difference in prediction error between two model-fitting algorithms. A naive implementation of the exhaustive nested cross-validation is computationally costly. Here, we address concerns regarding computational complexity by devising a computationally tractable closed-form expression for the proposed cross-validation estimator. Our study also investigates strategies aimed at enhancing statistical power within high-dimensional scenarios while controlling the Type I error rate. Through comprehensive numerical experiments, we demonstrate that Ridge-based methods using bias to measure uncertainty of CV estimates and adaptive hyperparameter selection provide the most reliable approach for high-dimensional predictive performance testing. To illustrate the practical utility of our method, we apply it to an RNA sequencing study and demonstrate its effectiveness in the context of biological data analysis.&lt;/p&gt;</description>
      <content:encoded><![CDATA[<p>It is crucial to assess the predictive performance of a model to establish its practicality and relevance in real-world scenarios, particularly for high-dimensional data analysis. Among data splitting or resampling methods, cross-validation (CV) is extensively used for several tasks such as estimating the prediction error, tuning the regularization parameter, and selecting the most suitable predictive model among competing alternatives. The K-fold cross-validation is a popular CV method but its limitation is that the risk estimates are highly dependent on the partitioning of the data (for training and testing). Here, the issues regarding the reproducibility of the K-fold CV estimator are demonstrated in hypothesis testing wherein different partitions lead to notably disparate conclusions. This study presents a novel predictive performance test and valid confidence intervals based on exhaustive nested cross-validation for determining the difference in prediction error between two model-fitting algorithms. A naive implementation of the exhaustive nested cross-validation is computationally costly. Here, we address concerns regarding computational complexity by devising a computationally tractable closed-form expression for the proposed cross-validation estimator. Our study also investigates strategies aimed at enhancing statistical power within high-dimensional scenarios while controlling the Type I error rate. Through comprehensive numerical experiments, we demonstrate that Ridge-based methods using bias to measure uncertainty of CV estimates and adaptive hyperparameter selection provide the most reliable approach for high-dimensional predictive performance testing. To illustrate the practical utility of our method, we apply it to an RNA sequencing study and demonstrate its effectiveness in the context of biological data analysis.</p>
]]></content:encoded>
    </item>
    
    <item>
      <title>Lead Statistician, Clinical Proteomics for Cancer Initiative: a proteomics-based discovery of non-small cell lung cancer biomarkers and drug targets in the Philippines</title>
      <link></link>
      <pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate>
      
      <guid></guid>
      <description></description>
      <content:encoded><![CDATA[]]></content:encoded>
    </item>
    
    <item>
      <title>Lead Statistician, Biosurveillance through Whole Genome Sequencing of SARS-CoV-2 in the Philippines</title>
      <link></link>
      <pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate>
      
      <guid></guid>
      <description></description>
      <content:encoded><![CDATA[]]></content:encoded>
    </item>
    
    <item>
      <title>Classification of Congenital Hypothyroidism using Artificial Neural Networks</title>
      <link></link>
      <pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate>
      
      <guid></guid>
      <description>&lt;p&gt;The Newborn Screening Reference Center (NSRC) of the National Health Institute in University of the Philippines Manila collects measurements from five attributes to determine whether Congenital Hypothyroidism (CH) is present in a neonate. Detecting the CH cases is a major concern of medical practitioners because it provides richer information than the healthy ones. However, because of the rarity of this metabolic condition, existing classification algorithms oftentimes misclassify a newborn as &amp;ldquo;normal&amp;rdquo; even if it is not. This paper investigates the efficiency of Self-Organizing Kohonen Maps (SOM), a type of artificial neural network. Though it is a visualization and clustering tool, the researchers want to probe on its ability to detect outliers and properly classify a newborn as normal or not by coming up with a statistically computed threshold value. Instead of working directly with the original attributes of the data, a reduced set of SOM prototypes is utilized to represent the data in a space of smaller dimension, seeking to preserve the probability distribution and topology of the input space. Results showed a misclassification rate of 13.5%. Though it is found to be slightly less superior to the existing classification rules, the proposed methodology was able to address the problem of finding a statistical threshold value. Also, the methodology verifies that age has a major effect on misclassifying &amp;ldquo;Normal&amp;rdquo; as &amp;ldquo;Abnormal&amp;rdquo; since postponement of newborn screening to a later age causes the quantization error to boost drastically, hence, easily exceeding the value of the first decision threshold.&lt;/p&gt;</description>
      <content:encoded><![CDATA[<p>The Newborn Screening Reference Center (NSRC) of the National Health Institute in University of the Philippines Manila collects measurements from five attributes to determine whether Congenital Hypothyroidism (CH) is present in a neonate. Detecting the CH cases is a major concern of medical practitioners because it provides richer information than the healthy ones. However, because of the rarity of this metabolic condition, existing classification algorithms oftentimes misclassify a newborn as &ldquo;normal&rdquo; even if it is not. This paper investigates the efficiency of Self-Organizing Kohonen Maps (SOM), a type of artificial neural network. Though it is a visualization and clustering tool, the researchers want to probe on its ability to detect outliers and properly classify a newborn as normal or not by coming up with a statistically computed threshold value. Instead of working directly with the original attributes of the data, a reduced set of SOM prototypes is utilized to represent the data in a space of smaller dimension, seeking to preserve the probability distribution and topology of the input space. Results showed a misclassification rate of 13.5%. Though it is found to be slightly less superior to the existing classification rules, the proposed methodology was able to address the problem of finding a statistical threshold value. Also, the methodology verifies that age has a major effect on misclassifying &ldquo;Normal&rdquo; as &ldquo;Abnormal&rdquo; since postponement of newborn screening to a later age causes the quantization error to boost drastically, hence, easily exceeding the value of the first decision threshold.</p>
]]></content:encoded>
    </item>
    
    <item>
      <title>Classification of Congenital Hypothyroidism in Newborn Screening Using Self-Organizing Maps</title>
      <link></link>
      <pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate>
      
      <guid></guid>
      <description>&lt;p&gt;Each day, the Newborn Screening Reference Center (NSRC) of the National Health Institute in University of the Philippines Manila collects measurements from five attributes to determine whether Congenital Hypothyroidism (CH) is present in a neonate. Detecting the CH cases is a major concern of medical practitioners because it provides richer information than the healthy ones. However, because of the rarity of this metabolic condition, existing classification algorithms oftentimes misclassify a newborn as &amp;ldquo;normal&amp;rdquo; even if it is not. This paper investigates the efficiency of Self-Organizing Kohonen Maps (SOM), a type of artificial neural network. Though it is widely known as a tool for visualization and clustering, the researchers want to probe on its ability as a tool for classification, particularly in detecting outliers. Results show that a lower misclassification rate yields from a self-organizing map with higher learning rate and larger training sample size. A bootstrap estimate of the variability of the misclassification error of roughly around 5% is also obtained. The misclassification error rate is lower when the original validation sample is used, compared to the average misclassification error rate computed from the bootstrap validation samples. Particularly, for a learning rate of 0.8 and a ratio of 2:1 training to validation sample, a 2.04% misclassification against 7.93% misclassification with 4.86% standard deviation is observed.&lt;/p&gt;</description>
      <content:encoded><![CDATA[<p>Each day, the Newborn Screening Reference Center (NSRC) of the National Health Institute in University of the Philippines Manila collects measurements from five attributes to determine whether Congenital Hypothyroidism (CH) is present in a neonate. Detecting the CH cases is a major concern of medical practitioners because it provides richer information than the healthy ones. However, because of the rarity of this metabolic condition, existing classification algorithms oftentimes misclassify a newborn as &ldquo;normal&rdquo; even if it is not. This paper investigates the efficiency of Self-Organizing Kohonen Maps (SOM), a type of artificial neural network. Though it is widely known as a tool for visualization and clustering, the researchers want to probe on its ability as a tool for classification, particularly in detecting outliers. Results show that a lower misclassification rate yields from a self-organizing map with higher learning rate and larger training sample size. A bootstrap estimate of the variability of the misclassification error of roughly around 5% is also obtained. The misclassification error rate is lower when the original validation sample is used, compared to the average misclassification error rate computed from the bootstrap validation samples. Particularly, for a learning rate of 0.8 and a ratio of 2:1 training to validation sample, a 2.04% misclassification against 7.93% misclassification with 4.86% standard deviation is observed.</p>
]]></content:encoded>
    </item>
    
    <item>
      <title>Project Leader, Sustainable Development Goals Research for Disability Registration Framework</title>
      <link></link>
      <pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate>
      
      <guid></guid>
      <description></description>
      <content:encoded><![CDATA[]]></content:encoded>
    </item>
    
    <item>
      <title>Classification of Multivariate Functional Data with an Application to ADHD fMRI Data</title>
      <link></link>
      <pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate>
      
      <guid></guid>
      <description>&lt;p&gt;The classification of resting-state functional magnetic resonance imaging (rs-fMRI) data presents unique challenges in the detection and diagnosis of neuropsychiatric disorders such as Attention-Deficit/Hyperactivity Disorder (ADHD). Traditional classification approaches often prove inadequate when handling complex spatiotemporal patterns and high-dimensional fMRI data, particularly when significant variations exist both between and within diagnostic groups. To address these limitations, we introduce a novel classification framework that integrates three complementary analytical components: elastic registration for curve alignment, geometric curve length computation for capturing signal variability, and sparse principal component analysis for dimensionality reduction. Extensive simulation studies show that our proposed method significantly outperforms existing approaches, especially in scenarios where groups exhibit distinct variation patterns rather than mean differences in their functional curves. When applied to the ADHD-200 dataset, our method achieves classification accuracy rates substantially exceeding conventional approaches. The proposed framework&amp;rsquo;s ability to capture subtle variability differences while maintaining computational efficiency makes it particularly valuable for biomarker discovery and clinical applications in neuropsychiatric research. Our approach&amp;rsquo;s focus on signal variability rather than mean activation patterns offers new insights into the dynamic nature of brain activity differences in ADHD and provides a promising foundation for analyzing other neurological conditions.&lt;/p&gt;</description>
      <content:encoded><![CDATA[<p>The classification of resting-state functional magnetic resonance imaging (rs-fMRI) data presents unique challenges in the detection and diagnosis of neuropsychiatric disorders such as Attention-Deficit/Hyperactivity Disorder (ADHD). Traditional classification approaches often prove inadequate when handling complex spatiotemporal patterns and high-dimensional fMRI data, particularly when significant variations exist both between and within diagnostic groups. To address these limitations, we introduce a novel classification framework that integrates three complementary analytical components: elastic registration for curve alignment, geometric curve length computation for capturing signal variability, and sparse principal component analysis for dimensionality reduction. Extensive simulation studies show that our proposed method significantly outperforms existing approaches, especially in scenarios where groups exhibit distinct variation patterns rather than mean differences in their functional curves. When applied to the ADHD-200 dataset, our method achieves classification accuracy rates substantially exceeding conventional approaches. The proposed framework&rsquo;s ability to capture subtle variability differences while maintaining computational efficiency makes it particularly valuable for biomarker discovery and clinical applications in neuropsychiatric research. Our approach&rsquo;s focus on signal variability rather than mean activation patterns offers new insights into the dynamic nature of brain activity differences in ADHD and provides a promising foundation for analyzing other neurological conditions.</p>
]]></content:encoded>
    </item>
    
    <item>
      <title>Analysis of Experiments with High Frequency Time Series Responses and the Implications for Power and Sample Size</title>
      <link></link>
      <pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate>
      
      <guid></guid>
      <description>&lt;p&gt;Given more accessible non-invasive measuring devices, experimental response can now be observed as high-dimensional and high-frequency time series. Amidst the complex dependence structure in the data analysis, sample size determination and power analysis remain to be the key thematic focus of statistical inference. The issue is confounded with the complexity of time lag structure and phase shift usually observed in a non-uniform but normal process typically present in medical imaging data. To address these issues in case-control studies, responses can be analyzed to obtain evidence of group differences through time series clustering based on dynamic time warping. The warping of multiple time series provides a flexible distance measure robust to time point concurrence. Time series clustering partitions experimental units into groups, enabling the computation of distances to measure effect size through sum of squares of pairwise distances in warped time series. Time series clustering provides an alternative to analysis of variance when experimental responses are high-frequency time series data. Kernel regression is formulated to link sample size, effect size, power of the test, and level of significance accounting for the structure of the data generating process of the time series responses. This provides a strategy for clinicians to optimize the power of the test that can be achieved with a minimal sample size for this experimental setup. Time series clustering method is able to differentiate case and control groups in the simulated data and in the ADHD-200 fMRI dataset. The distance measured between two or more groups of time series can be used to determine sample size for a target power.&lt;/p&gt;</description>
      <content:encoded><![CDATA[<p>Given more accessible non-invasive measuring devices, experimental response can now be observed as high-dimensional and high-frequency time series. Amidst the complex dependence structure in the data analysis, sample size determination and power analysis remain to be the key thematic focus of statistical inference. The issue is confounded with the complexity of time lag structure and phase shift usually observed in a non-uniform but normal process typically present in medical imaging data. To address these issues in case-control studies, responses can be analyzed to obtain evidence of group differences through time series clustering based on dynamic time warping. The warping of multiple time series provides a flexible distance measure robust to time point concurrence. Time series clustering partitions experimental units into groups, enabling the computation of distances to measure effect size through sum of squares of pairwise distances in warped time series. Time series clustering provides an alternative to analysis of variance when experimental responses are high-frequency time series data. Kernel regression is formulated to link sample size, effect size, power of the test, and level of significance accounting for the structure of the data generating process of the time series responses. This provides a strategy for clinicians to optimize the power of the test that can be achieved with a minimal sample size for this experimental setup. Time series clustering method is able to differentiate case and control groups in the simulated data and in the ADHD-200 fMRI dataset. The distance measured between two or more groups of time series can be used to determine sample size for a target power.</p>
]]></content:encoded>
    </item>
    
    <item>
      <title>Testing the Homogeneity of Risk Differences with Sparse Count Data</title>
      <link></link>
      <pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate>
      
      <guid></guid>
      <description>&lt;p&gt;In this paper, we consider testing the homogeneity of risk differences in independent binomial distributions especially when data are sparse. We point out some drawback of existing tests in either controlling a nominal size or obtaining powers through theoretical and numerical studies. The proposed test is designed to avoid the drawbacks of existing tests. We present the asymptotic null distribution and asymptotic power function for the proposed test. We also provide numerical studies including simulations and real data examples showing the proposed test has reliable results compared to existing testing procedures.&lt;/p&gt;</description>
      <content:encoded><![CDATA[<p>In this paper, we consider testing the homogeneity of risk differences in independent binomial distributions especially when data are sparse. We point out some drawback of existing tests in either controlling a nominal size or obtaining powers through theoretical and numerical studies. The proposed test is designed to avoid the drawbacks of existing tests. We present the asymptotic null distribution and asymptotic power function for the proposed test. We also provide numerical studies including simulations and real data examples showing the proposed test has reliable results compared to existing testing procedures.</p>
]]></content:encoded>
    </item>
    
    <item>
      <title>Nonparametric Modeling of Clustered Customer Survival Data</title>
      <link></link>
      <pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate>
      
      <guid></guid>
      <description>&lt;p&gt;We incorporate a random clustering effect into the nonparametric version of Cox Proportional Hazards model to characterize clustered survival data. The simulation studies provide evidence that clustered survival data can be better characterized through a nonparametric model. Predictive accuracy of the nonparametric model is affected by number of clusters and distribution of the random component accounting for clustering effect. As the functional form of the covariate departs from linearity, the nonparametric model is becoming more advantageous over the parametric counterpart. Finally, nonparametric is better than parametric model when data are highly heterogenous and/or there is misspecification error.&lt;/p&gt;</description>
      <content:encoded><![CDATA[<p>We incorporate a random clustering effect into the nonparametric version of Cox Proportional Hazards model to characterize clustered survival data. The simulation studies provide evidence that clustered survival data can be better characterized through a nonparametric model. Predictive accuracy of the nonparametric model is affected by number of clusters and distribution of the random component accounting for clustering effect. As the functional form of the covariate departs from linearity, the nonparametric model is becoming more advantageous over the parametric counterpart. Finally, nonparametric is better than parametric model when data are highly heterogenous and/or there is misspecification error.</p>
]]></content:encoded>
    </item>
    
    <item>
      <title>Modelling Clustered Survival Data with Cured Fraction</title>
      <link></link>
      <pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate>
      
      <guid></guid>
      <description>&lt;p&gt;In modelling lifetime data, standard parametric theory assumes that all observations will eventually experience the event of interest if they are monitored for a very long period. While every unit starts as susceptible to the event of interest, a fraction of observations may switch into a non-susceptible group. A mixture cured fraction model with covariates is modified to incorporate random clustering effect to characterize the switch mechanism. Simulation studies and telecommunications data show that cured fraction models with random clustering effect perform better than their parametric counterpart in terms of predictive ability.&lt;/p&gt;</description>
      <content:encoded><![CDATA[<p>In modelling lifetime data, standard parametric theory assumes that all observations will eventually experience the event of interest if they are monitored for a very long period. While every unit starts as susceptible to the event of interest, a fraction of observations may switch into a non-susceptible group. A mixture cured fraction model with covariates is modified to incorporate random clustering effect to characterize the switch mechanism. Simulation studies and telecommunications data show that cured fraction models with random clustering effect perform better than their parametric counterpart in terms of predictive ability.</p>
]]></content:encoded>
    </item>
    
  </channel>
</rss>
