Personal Website
Search Results
Search this site
67 results found with an empty search
- Understanding Poisson Distribution and Modeling in R
This week in my 602 course I learned about the Poisson distribution. The Poisson distribution is a fundamental tool in statistics for modeling count data — situations where the outcome is the number of times an event occurs in a fixed interval of time, space, or another dimension. This distribution assumes that events happen independently and at a constant average rate. When the mean (which is also the variance) increases, this approaches binomial distribution, and if it increases even more it approaches a normal distribution. Here is a photo of the French mathematician. 🙏 When to Use Poisson Models? Poisson models are used when: The outcome variable is a count (e.g., number of ER visits, crashes, phone calls). Events are rare and occur independently. The variance of the outcome is equal to its mean (equidispersion). Using "person-time" allows for each person's follow up time to contribute meaningfully. Below we see 18 person years total. Modeling Poisson in R Here is a basic example using R: # Simulated data set.seed(123) data <- data.frame( age = sample(20:60, 100, replace = TRUE), gender = sample(c("M", "F"), 100, replace = TRUE), exposure = runif(100, 1, 10) ) data$events <- rpois(100, lambda = 0.3 * data$exposure) # Fit Poisson regression model <- glm(events ~ age + gender + offset(log(exposure)), data = data, family = poisson) summary(model) This fits a Poisson regression with an offset to account for varying exposure time. Problem 1: Overdispersion A core assumption of Poisson regression is that the variance equals the mean. However, in real-world data, variance often exceeds the mean. This is called overdispersion. This could happen if the model is missing an interaction term, has some non linear covariates, or missing an important variable. Symptoms of Overdispersion: Residual deviance much greater than degrees of freedom Large standard errors Poor model fit Fixing Overdispersion: One common way is to use: Negative Binomial regression: Adds a dispersion parameter. library(MASS) nb_model <- glm.nb(events ~ age + gender + offset(log(exposure)), data = data) Problem 2: Zero Inflation Some count data have more zeros than expected under a Poisson model. This is common in biological, health, or social data. This leads to zero inflation. Consider the below red graph for the number of hospital visits in a year. There could be patients who never go see a doctor. There could be patients just happened to not go this year. The zeros are inflated. Solution: Zero-Inflated Models Zero-Inflated Poisson (ZIP) and Zero-Inflated Negative Binomial (ZINB) models use a two-part process: A logistic model to predict excess zeros (who are always zero, versus happened to be zero) A count model for non-zero counts (remove the predicted always zeros and refit poisson) library(pscl) zip_model <- zeroinfl(events ~ age + gender | 1, data = data, dist = "poisson") Summary Poisson regression is powerful for modeling rates and counts, incidence rate in epidemiology is a bis area for this, but real-world data could violate its assumptions. Be on the lookout for overdispersion and zero inflation, and use alternative models like Negative Binomial or Zero-Inflated Poisson when needed. I hope you find this post helpful! Let's keep learning!
- Pregnancy & Healthy Eating
👶🥗 Pregnancy & Healthy Eating — We’ve Got Work To Do A new NIH study reveals a surprising truth: most people in the U.S. don’t follow healthy diets during or after pregnancy — despite knowing how important nutrition is for parent and baby health. 📉 Using data from 10,000+ people, the study shows: Very few pregnant or postpartum individuals meet national dietary guidelines. People with lower income and education are especially affected. Gaps in nutrition persist even months after giving birth. 🧠 Why it matters:What we eat in pregnancy isn’t just about today — it lays the foundation for long-term health outcomes for both parent and child. Poor diets can increase risks of: Low birth weight Gestational diabetes Obesity and heart disease later in life 🛍️ Access, affordability, and culturally relevant food support are critical. It’s not about blaming individual choices — it’s about building systems that make healthy options accessible and sustainable . Let’s make maternal health holistic — support policies that bring nutritious food, education, and equity into every community. 📎 Read the full update from NIH here: Healthy eating not common during and after pregnancy
- Healthy Aging
🧠🥦 Want to age well? Your midlife diet really matters -- big time. A new long-term study in Nature Medicine followed 100,000+ people for 30 years and found that people who stuck to healthy diets in midlife were more likely to age in good health. ✅ The best-performing diet was the Alternative Healthy Eating Index (AHEI) – rich in fruits, veggies, whole grains, nuts, and healthy fats, with less red meat, processed foods, and sugar. 📊 People in the healthiest diet group had up to 86% better odds of aging well — meaning they were more likely to avoid chronic diseases, maintain physical and mental health, and live past 70. 🍩 On the flip side? Diets high in ultra processed foods (think chips, soda, fast food) were linked to significantly lower chances of aging in good shape. What can we do? If there was a mark on the food packages showing how ultra processed they are, we could have made more informed decisions. Not only considering the ingredients though, both ingredients and the processing that the food is going through. This study shows it’s not just about living longer — it’s about staying sharp, mobile, and happy while doing it. 👵🏽👴🏽💚 So, the earlier you start eating smart, the better your shot at a healthier, brighter future. ✅ Foods that boost healthy aging: Vegetables (+0.86) Fruits Whole Grains Nuts & Legumes Healthy Fats ❌ Foods that hurt healthy aging: Red/Processed Meats Sugary Drinks Ultra-Processed Foods (worst impact) Here’s a visual summary (from the paper) of the best and worst foods for healthy aging👇 Do you know some people who aged in an extremely healthy way? 🔗 Read the full study: https://www.nature.com/articles/s41591-025-03570-5
- Crusty Bread and Joint Pain: Is There a Connection?
Do you suffer from joint pain? Do you love crusty, crispy breads, toasted breads, and chips? These could be linked! In Turkey, bread is a staple at every meal. To be considered a good homemaker, you’d better get fresh, crispy-crusted bread from the bakery every single day. Inside this golden loaf? Pretty much air—you can break it in half with your hands and stuff it with tomato slices, feta cheese, and olives for a quick teenage meal. When I was a kid, I loved nibbling on the crust on my way home from the bakery. By the time I got home, my mom would yell, “What the hell happened to this bread?!” 😂 Joint Pain in Turkey According to this article , 15% of adults in Turkey experience knee pain ( 26,4% in women and 6,2% in men) . Many seek medical treatments, from stem cell injections into joints to alternative remedies like: Chiropractic adjustments Acupuncture Cupping therapy Leech treatments (Yes, really!) But what if this aging-related joint pain isn’t just about wear and tear? What if it’s related to diet, specifically to AGEs in crispy, crusted breads? What Are AGEs? And Why Should You Care? AGEs (Advanced Glycation End-products) are compounds formed when sugars react with proteins at high temperatures. Think: 🔥 Honey roasted Peanuts 🔥Crusty bread 🔥 Toasted foods 🔥 Chips and croutons. Scientists have been feeding experimental models bread crust to test its effects on bone and joint health. The results? Not great. 📉 Bread crust reduced bone stiffness by 50%. 📉 Collagen fibers in joints were damaged. 📉 The tail (a flexible structure in experimental models) became stiff. Here are some direct research findings: 📌 “Food-derived AGEs may alter tendon properties and lead to tendon injuries.” ( PubMed ) 📌 “Dietary AGEs can cause cross-links in tendons, affecting tissue compliance and tendon health.” ( PubMed ) 📌 “Consumption of Maillard Reaction Products (MRPs) leads to lower bone size and changes in mechanical behavior.” ( PubMed ) How This Changed My Thinking My husband recently started seeing a chiropractor for neck pain that causes uncomfortable popping sounds in his ears. I suggested he see an internist, but he opted for massages first. The chiropractor told him his body and head are slightly tilted to one side. This made me reevaluate my own posture. I always carry my computer tote on my right shoulder—is it tilting my body too? 🤔 And then, after reading the bread crust studies, I looked at our high-capacity toaster with suspicion. I even asked my husband to reconsider his chip addiction. Croutons? Maybe not the best addition to my salad after all. What I’m Doing Now I can’t stop aging or aging-related discomfort, but here’s what I’m taking every day to support my body: (Does not constitute medical advice) ✅ Halal multivitamin ✅ Halal vitamin D (could be once a week) ✅ Creatine gummies (could be once a week) ✅ Probiotic gummies ✅ Garlique (garlic extract) ✅ Omega 3 (could be once week) ✅ Fiber pills I’ve also started opting for soft, no-crust breads whenever possible. What About You? Do you have memories with crispy bread? Do you or someone you know suffer from joint pain? Have they tried specific remedies—like boiling cherry stems or other traditional treatments? Drop a comment! Let’s discuss. 👇😊
- Experiments and Sanity
Sometimes our experiments get very complicated and there are some tricks that can help prepare better. Today I want to share with you what helps me navigate high-stress experiment days like these. 1️⃣ Plan, Plan, Plan Some days, I dedicate time just to planning: mapping out protocols, gathering supplies, and mentally walking through the entire experiment to anticipate challenges. This saves me from last-minute surprises. 2️⃣ Keep Master Protocols Every time I learn a new experiment, I start with someone’s written protocol and then refine it. I add details like exact locations of materials, catalog numbers, and whether incubation for 1 hour vs. 2 hours really matters. The little details can make a big difference. 3️⃣ Check Inventory (Before It’s Too Late) Ever started an experiment only to realize you’re out of a critical reagent? I’ve had my fair share of running from lab to lab, hunting for 5 µL of something I thought I had. Now, I check inventory before starting. Lesson learned. 4️⃣ Work in Teams For big experiments—having a helper is a game changer. Team-tagging with co-authors or lab mates makes everything more efficient (and keeps you sane). 5️⃣ Stay Clean & Organized Decluttering, decontaminating, and staying as clean as possible actually makes a difference. If nothing else, it makes you feel better. Trust me. 6️⃣ Reward Yourself Difficult experiment ahead? I mentally prepare by setting up a reward system. My personal trick? Moving beads from one container to another as a little victory tracker. Sounds silly, but it works! What’s one experiment or protocol you struggled with but eventually mastered? Any tricks that made it easier?
- Hitch hiking Through the Subconscious: A Journey with t-SNE and Logistic Regression
Yesterday in My Dream… I told the driver of a car that I hitchhiked that I was doing a PhD in Computer Science. (What kind of subconscious is this?) The road was dark, endless. I couldn't see the end. So, I decided to hitchhike. Then, I saw a mini Turtle car—light green, safe-looking—so I trusted it. I held up my arm, and inside was another student. The driver’s seat? In the back. 🚗💨 Through the darkness, we drove, and we arrived in the city. Antalya, apparently. I asked him, “Which way do you need to go?” He said, “I need to turn right.” So, I got off at the intersection. Then, I noticed I forgot my belongings in the car. My phone and a red paper bag were left behind. I went back and found them neatly placed on the sidewalk. How very considerate of my subconscious driver. Logistic Regression in My Dreams?! I had been studying logistic regression all weekend. Apparently, my brain decided this meant I was doing a PhD in Computer Science. 💭 This is why we don’t cram for exams, people. To recover, we watched Güldür Güldür to reset our happiness levels. One quote stood out: “If your conscious is a mirror, then your subconscious is the cover behind the mirror.” (Aynanın ardındaki sır.) Profound. You may want to watch how Turkish comedians view Trump as well. :) How to Become a Millionair Using LinkedIn Overnight 1️⃣ Take a deep breath in. 2️⃣ Drop the “e” from millionaire. That’s it. 💸😂 P.S. I should acknowledge my inspiration: A black van driving in front of me on I-64 West today. www.millionair.com 🤦♀️ P.P.S. My mental state was perfectly prepped for this joke after listening to this song 20 times: 🎵 P.P.P.S. Also restarted my Vitamin D pills (5000 IU/day), because, you know, forced richness starts with forced serotonin. 🌞💊 From Forced Happiness to Forced Education: Statistics Edition My next goal? A forced education on statistics formulas. Time to drive pluripotency of my grey region neurons to wake up, proliferate, and make new connections. I even suggested an algorithm to ChatGPT yesterday. It said, “This already exists. It’s called UMAP.” 🤡 This was as funny as ChatGPT asking me to verify that I am a human. BRO, WHAT DO YOU MEAN? Final Thoughts: Water Your Plants 🌱 Enjoy your Tuesday. Take your vitamins. And remember: Water your plants, or they’ll use t-SNE to dimensionality reduce themselves into the void. 🌱
- Can We Prevent Cancer? 🛡️🧬
In my last post, we talked about what happens when a cell becomes a cancer cell. Today, let’s talk about environmental and preventable risk factors that contribute to cancer.. We know that genetics plays a role, and some cancers are inherited. But beyond genetics, there are layers of risk factors that we can control. 🔬 Have You Heard of Epigenetics? Epigenetics refers to factors that influence gene expression without changing the actual DNA sequence. It can protect us or increase our risk of disease. Think of your DNA as a thread wrapped around small balls called nucleosomes. The accessibility of DNA is controlled by epigenetic modifications. 🧬 If DNA is tightly packed (heterochromatin), it’s less accessible. 🛑 If DNA is looser (euchromatin), certain genes—including disease-causing ones—may become active. 🏠 How Does the Environment Influence Cancer Risk? Your lifestyle and surroundings can modify chromatin structure, making harmful genes more or less active. For example: ✅ If you inherit harmful mutations but they stay silent, your risk is lower. ❌ If your environment triggers those genes to become active, your risk increases. ⚠️ Key Environmental Risk Factors for Cancer: 1️⃣ Unhealthy Diet & Chronic Inflammation – Obesity, ultra-processed foods, and inflammatory triggers increase cancer risk. 🥤🍔 2️⃣ Tobacco smoking, air pollution, asbestos exposure. 3️⃣ Chemical Exposure – Certain cosmetics, hair products, plastics, and banned chemicals can disrupt cellular health. 🧴☠️ 4️⃣ Disrupted sleep, impacting circadian rhythm. 5️⃣ Ultraviolet light exposure is not good for skin. 6️⃣ Tattoos? – A recent study suggests that tattoo inks may pose risks. 🖋️ 🏃 What Can We Do? While we can’t control our genes, we can lower our risk by: 🥗 Eating whole, natural foods and reducing inflammatory triggers. 🚿 Minimizing chemical exposure—be mindful of cosmetics & household products. 🤔 Considering the long-term effects before making lifestyle choices (e.g., tattoos). 🔄 What Else Comes to Mind? What other preventative measures do you think we should talk about? Drop your thoughts below! ⬇️ We will talk more about ultra processed foods next.
- 🍖 The Hidden Danger of AGEs in Food
We often talk about ultra-processed foods, but how they’re cooked matters too! 🍗🔥 When foods undergo high heat processing—like frying, roasting, or grilling—sugars attach to proteins and fats, forming harmful compounds called Advanced Glycation End Products (AGEs). 🍯🥜 What Foods Are High in AGEs? 🟢 Low AGE foods: Fresh fruits 🍏, vegetables 🥕, dairy 🥛, and whole grains 🌾. 🔴 High AGE foods: Fried chicken 🍗, grilled meats 🥩, honey-roasted nuts 🥜, and processed snacks 🏭. ⚠️ Why Should We Care About AGEs? AGEs are linked to: 🚨 Premature aging and organ damage 🏥 🔥 Chronic inflammation and metabolic disorders 🔥 🦠 Cancer growth & metastasis Can We Reduce AGEs in Our Diet? ✅ Cook at lower temperatures (steaming, boiling, slow-cooking). ✅ Swap fried & grilled for baked or raw options. ✅ Eat more whole, unprocessed foods. This research raises big questions for public health. Should we be more mindful of not just what we eat, but how we cook it? Let’s discuss in the comments! ⬇️
- 🥐 Ultra-Processed Foods – How Much Are We Really Eating?
Did you know? A recent NIH study found that ultra-processed foods (UPFs) make up 70% of the American diet! 🍔🥤 🔬 This study, led by Dr. Kevin Hall at NIH Bethesda, had participants live in a controlled setting, tracking their food intake, body measurements, metabolism, and even the air they breathed out! 🫁🔥 (I will put the link in the comments) 🍞 What Counts as Ultra-Processed? Even foods we consider “healthy”—like sliced bread—often contain additives and preservatives that push them into the ultra-processed category. 🏭❌ A landmark study used AI & machine learning to analyze food labels in grocery stores. Results? A staggering 73% of U.S. food supply is ultra-processed! 🤯 🔹 Example: A raw onion = minimally processed 🧅 ✅, but frozen onion rings = ultra-processed 🏭❌. The processing changes its natural components more than 10-fold! ⚠️ Why Does This Matter? Eating UPFs makes it harder for our bodies to regulate hunger, leading to overeating, metabolic diseases, and increased health risks. The big question: Should governments regulate UPFs the same way they regulate tobacco & alcohol? 🤔 Drop your thoughts below! ⬇️
- 🧬 What Do Your Cells Do in a Day?
Ever wonder what your cells are up to? 🤔 The first lecture in Molecular Cell Biology always starts with just how small cells are. A typical cell is 10 µm by 15 µm—so tiny that I used to ask my students: 💡 How many cells do you think fit under your pinky nail? If you think of it as a 2D surface, the answer is ~670,000! 😲 That’s a LOT of cells—and that’s just your pinky! Now imagine your entire body! 🛠️ What Do Cells Actually Do All Day? It depends on where they are! Cells are masters of teamwork 🏗️. They’ve evolved to function as part of a community, sharing work perfectly and communicating at almost miraculous levels. Cells are alive: ✅ They move, eat, and produce waste 💩 ✅ They send & receive signals 📡 to talk to each other ✅ They adjust to their environment 🏗️ by changing their cytoskeleton, gene expression, or even dividing into two. I used to joke in class: 💬 “What is a Molecular Cell Biology teacher? Just a pile of cells… talking about cells.” 😂 🧬 Not All Cells Are the Same! Recent single-cell sequencing technologies 🧪 have confirmed something fascinating: 🌀 Each cell is unique—no two are exactly alike! This diversity comes from layers of regulation, allowing infinite combinations of gene expression and function. Let’s take a sneak peek at a few specialized cells: 🍽️ Intestinal Cells: The Ultimate Nutrient Transporters Your intestinal epithelial cells work hard to absorb glucose & nutrients from your food. 🥞 🔹 If these cells stopped working, everything you ate would just pass through as waste! 😱 🔹 Even if there's just ONE glucose molecule, these cells will grab it and push it into the bloodstream using the Na+/Glucose Symporter and the Na+/K+ pump. (Okay, this part is getting technical! I will post a link to dive deeper if you’re curious!) ⚡ Excitable Cells: The Body’s Electricians! Did you know? 💡Your neurons and muscle cells experience electric fields of 200,000 volts per cm (V/m) across their membranes—nearly 100,000 times stronger than high-voltage power lines! ⚡😵 So how do they survive this? Think of them as skilled dancers 💃—they perform a Mexican wave-like motion to pass messages along, using action potentials, which again involves ion channels and pumps. ⚡ Neurons send signals → Muscle cells contract → You move your finger! (Muscle contraction is a BIG topic—if you want to see how it works, check out this video in the comments! 🎥) 🔬 So Many More Amazing Cells! We haven’t even talked about: 🩸 Blood cells 🏋️ Smooth muscle cells 🛑 Adipose (fat) cells I’ll save those for another post! 😉 🎯 The Big Picture Your cells are busy—working 24/7 so that you can thrive, learn, grow, and experience life. Take a moment to thank the tiny beings that make you who you are. 🫶 Thank you, little me! 🧬✨
- How does tSNE work?
A Deep Dive into Stochastic Neighbor Embedding Introduction In high-dimensional data visualization, t-Distributed Stochastic Neighbor Embedding (t-SNE) has become one of the most popular techniques. It allows us to map complex datasets into two or three dimensions while preserving local structures. But how does it work? In this post, we will break down t-SNE step by step, drawing insights from its mathematical foundations and practical applications. 1. What is t-SNE Trying to Solve? When we work with high-dimensional data (e.g., images, gene expression data, or word embeddings), understanding its structure becomes difficult. We need a way to: ✅ Reduce the data to 2D or 3D for visualization. ✅ Preserve clusters and local relationships. ✅ Avoid distortions caused by simple linear transformations (like PCA). Unlike PCA, which is a linear method, t-SNE is non-linear and focuses on preserving local neighborhoods. This makes it ideal for detecting clusters and substructures in data. 2. What Does "Stochastic" Mean in t-SNE? "Stochastic" means random but with structure . In t-SNE: We start by randomly positioning the data points in low dimensions. We then iteratively adjust their positions, using probabilities to match their high-dimensional relationships. This randomness helps avoid bad solutions and ensures better local clustering. Because of this, t-SNE’s output can vary slightly across runs—unlike PCA, which always gives the same result: tSNE with seed 42 tSNE with seed 100 3. The Key Idea Behind t-SNE At its core, t-SNE works by modeling similarities between points in both high-dimensional and low-dimensional spaces: Step 1: Compute Pairwise Similarities in High-Dimensional Space Each data point gets a probability score for how similar it is to other points. This is done using a Gaussian distribution centered at each point. Similar points have high probability, while distant points have low probability. Step 2: Compute Pairwise Similarities in Low-Dimensional Space Instead of a Gaussian, we use a t-distribution (which has heavier tails to avoid overcrowding). The goal is to make sure that the low-dimensional probabilities match the high-dimensional ones as closely as possible. Step 3: Minimize the Difference (KL Divergence) The "error" function t-SNE minimizes is called Kullback-Leibler (KL) Divergence . This function tells us how different two probability distributions are. Using gradient descent, t-SNE moves points in 2D space until the low-dimensional structure resembles the high-dimensional one . 4. Why Does t-SNE Sometimes Look Different Every Time? Since t-SNE is stochastic, each run may produce a slightly different layout. This happens because: The algorithm starts with a random initialization of points. The optimization process can get stuck in different local minima. Different perplexity values (a hyperparameter) can change how local/global structure is balanced. 🔹 Solution: Run t-SNE multiple times and look for stable patterns! You can also set a seed for reproducibility.
- Famous Swiss Roll example
Understanding Dimension Reduction: PCA vs. t-SNE vs. UMAP with the Swiss Roll When dealing with high-dimensional data, it is often challenging to visualize and interpret patterns. Dimensionality reduction helps us project complex datasets into fewer dimensions while preserving meaningful structure. In this post, we will explore two popular dimensionality reduction techniques—Principal Component Analysis (PCA) and t-Distributed Stochastic Neighbor Embedding (t-SNE)—using a classic example: the Swiss Roll dataset. The Swiss Roll Dataset The Swiss Roll is a synthetic dataset where points are arranged on a curved 3D surface, forming a spiral-like shape. The challenge in dimensionality reduction is to flatten this 3D spiral while preserving important relationships between points. In this figure, each point is color-coded based on its position along the spiral. Ideally, after dimensionality reduction, points with similar colors should remain close together. PCA: Preserving Global Structure Principal Component Analysis (PCA) is a linear dimensionality reduction method. It works by identifying the directions (principal components) that capture the most variance in the data and projecting the dataset onto these components. When PCA is applied to the Swiss Roll, we obtain the following projection: Key Observations: ✅ PCA successfully collapses the Swiss Roll into two dimensions. ✅ It preserves the global structure, meaning the overall spiral shape is still visible. 🚨 However, PCA distorts local relationships—some points that were far apart in 3D may appear close in 2D. 📌 PCA is useful when the data follows a linear structure and we need a fast, interpretable projection. t-SNE: Preserving Local Structure t-Distributed Stochastic Neighbor Embedding (t-SNE) is a non-linear dimensionality reduction technique that focuses on preserving local relationships in the data. Instead of a linear projection, t-SNE maps points into 2D space while trying to maintain the pairwise similarities between points. Applying t-SNE to the Swiss Roll results in: tSNE projection Key Observations: ✅ t-SNE successfully unrolls the Swiss Roll by grouping similar points together. ✅ It preserves local neighborhoods—points that were close in 3D remain close in 2D. 🚨 However, t-SNE does not preserve the global spiral shape; instead, it breaks the roll into clusters. 📌 t-SNE is great for discovering clusters in high-dimensional data but does not maintain global structure. Comparison: PCA vs. t-SNE Feature PCA (Linear) t-SNE (Non-Linear) Preserves Global Structure? ✅ Yes ❌ No Preserves Local Structure? ❌ No ✅ Yes Works Well for Curved Data? ❌ No ✅ Yes Good for Cluster Discovery? ❌ No ✅ Yes Computationally Efficient? ✅ Fast ❌ Slow If you need a quick, structured projection, PCA is a great choice. If you need to find clusters and preserve neighborhood relationships, t-SNE is better. Both PCA and t-SNE are powerful tools for reducing dimensionality, but they serve different purposes. PCA is best when global structure matters, while t-SNE is useful for local clustering. Understanding these differences is key when visualizing and interpreting high-dimensional datasets. I am still debating, if I understood the whole concept. But this is essentially it. :) What about UMAP? At this point, I questioned, don't we need an algorithm that can preserve the spiral shape better? The answer is UMAP. We can argue this is still not the perfect spiral shape but better than the tSNE's broken apart pieces. UMAP Prioritizes Local Structure UMAP is designed to preserve local distances more than global distances. It maintained the order of the Swiss Roll but didn’t force it into a perfectly smooth spiral. Why UMAP Might Be the Best Choice for You Feature PCA t-SNE UMAP Preserves Global Structure? ✅ Yes ❌ No ✅ Yes Preserves Local Structure? ❌ No ✅ Yes ✅ Yes Distributes Points Naturally? ❌ No (line) ❌ No (overclusters) ✅ Yes (balanced) Fast for Large Data? ✅ Yes ❌ No (slow) ✅ Yes (faster) Have you used any of these algorithms? How was your experience?











