Highlighting the Molecular Bridge Between Insulin Resistance and Alzheimer's Disease¶
A data story built using NIH Common Fund Data Ecosystem (CFDE) resources¶
The Question: Can we use publicly available NIH biomedical datasets to identify the molecular mechanisms linking peripheral insulin resistance to Alzheimer's disease risk and other risk factors?
Approach: Using CFDE REVEAL's mechanism discovery tool and GTEx gene expression data, we traced a biological pathway from metabolic dysfunction in peripheral tissues to neurodegeneration in the brain.
import pandas as pd
import matplotlib.pyplot as plt
import matplotlib.patches as mpatches
import seaborn as sns
import warnings
warnings.filterwarnings('ignore')
sns.set_theme(style="whitegrid")
plt.rcParams['figure.dpi'] = 150
data = {
'Gene': ['APOE', 'IRS1', 'FOX03', 'FOX01', 'IRS2', 'SLC2A2', 'IGF1R', 'INSR', 'HMGA2'],
'Combined': [8.48, 8.33, 5.81, 5.64, 4.67, 4.55, 4.23, 3.79, 3.58],
'GWAS': [2.48, 4.05, 3.90, 1.17, 0.25, 1.92, 0.39, 1.55, 3.64],
'Functional': [6.00, 4.30, 1.91, 4.47, 4.43, 2.62, 3.86, 2.24, -0.07]
}
df = pd.DataFrame(data)
df
| Gene | Combined | GWAS | Functional | |
|---|---|---|---|---|
| 0 | APOE | 8.48 | 2.48 | 6.00 |
| 1 | IRS1 | 8.33 | 4.05 | 4.30 |
| 2 | FOX03 | 5.81 | 3.90 | 1.91 |
| 3 | FOX01 | 5.64 | 1.17 | 4.47 |
| 4 | IRS2 | 4.67 | 0.25 | 4.43 |
| 5 | SLC2A2 | 4.55 | 1.92 | 2.62 |
| 6 | IGF1R | 4.23 | 0.39 | 3.86 |
| 7 | INSR | 3.79 | 1.55 | 2.24 |
| 8 | HMGA2 | 3.58 | 3.64 | -0.07 |
Hypothesis 1: The Insulin Signaling Pathway¶
The first thing CFDE REVEAL showed was a classic insulin signaling chain INSR/IGF1R → IRS → PI3K/AKT → FOXO as a likely route connecting insulin resistance in the body to Alzheimer's risk in the brain.
hen this pathway breaks down in metabolic tissues, it sets off a chain reaction: glucose transport gets disrupted, key transcription factors stop working properly, and downstream processes tied to APOE start to shift.
What stood out: APOE and IRS1 ranked highest of all candidate genes, and unlike some others, they're backed by both population genetics data and lab-based experimental evidence not just one or the other.
# Hypothesis 2 data from CFDE REVEAL - the lipid transport pathway
# PPARG-APOE/PLTP adipose lipidation axis
data2 = {
'Gene': ['PPARG', 'APOE', 'IRS1', 'ADIPOQ', 'LPL', 'APOA1', 'PLTP', 'INSR'],
'Combined': [8.73, 8.48, 8.33, 6.75, 6.74, 4.84, 3.99, 3.79],
'GWAS': [3.54, 2.48, 4.05, 4.36, 3.87, 1.77, 2.68, 1.55],
'Functional': [5.20, 6.00, 4.30, 2.40, 2.87, 3.07, 1.31, 2.24]
}
df2 = pd.DataFrame(data2)
df2
| Gene | Combined | GWAS | Functional | |
|---|---|---|---|---|
| 0 | PPARG | 8.73 | 3.54 | 5.20 |
| 1 | APOE | 8.48 | 2.48 | 6.00 |
| 2 | IRS1 | 8.33 | 4.05 | 4.30 |
| 3 | ADIPOQ | 6.75 | 4.36 | 2.40 |
| 4 | LPL | 6.74 | 3.87 | 2.87 |
| 5 | APOA1 | 4.84 | 1.77 | 3.07 |
| 6 | PLTP | 3.99 | 2.68 | 1.31 |
| 7 | INSR | 3.79 | 1.55 | 2.24 |
Hypothesis 2: The Lipid Transport Pathway¶
CFDE REVEAL also showed a second, completely independent mechanism this one centered on fat tissue rather than direct insulin signaling. The key point here is that PPARG, a master regulator of fat cell biology, gets disrupted in insulin-resistant adipose tissue. That disruption throws off how APOE gets packaged and transported, which can mess with the brain's ability to clear amyloid beta a key factor of Alzheimer's pathology.
What stood out: PPARG and APOE top this list, with IRS1 close behind.
df2_sorted = df2.sort_values('Combined', ascending=True)
fig, ax = plt.subplots(figsize=(10, 6))
colors2 = ['#e07b54' if row['GWAS'] > row['Functional'] else '#5b8db8'
for _, row in df2_sorted.iterrows()]
ax.barh(df2_sorted['Gene'], df2_sorted['Combined'], color=colors2, edgecolor='white', height=0.6)
for i, (val, gene) in enumerate(zip(df2_sorted['Combined'], df2_sorted['Gene'])):
ax.text(val + 0.1, i, f'{val}', va='center', fontsize=9)
ax.set_xlabel('Relevance Score', fontsize=11)
ax.set_title('Candidate Genes Linking Adipose Lipidation to Alzheimer\'s Disease Risk\nSource: CFDE REVEAL',
fontsize=12, fontweight='bold')
gwas_patch = mpatches.Patch(color='#e07b54', label='Population evidence dominates (GWAS)')
func_patch = mpatches.Patch(color='#5b8db8', label='Experimental evidence dominates (Functional)')
ax.legend(handles=[gwas_patch, func_patch], loc='lower right', fontsize=9)
plt.tight_layout()
plt.savefig('gene_scores_h2.png', dpi=150, bbox_inches='tight')
plt.show()
Where the Two Pathways Converge¶
Here's the interesting part. From a single search, CFDE REVEAL surfaced two distinct mechanistic hypotheses one centered on insulin signaling, the other on fat tissue biology. Different mechanisms, different gene sets for the most part, but they both point back to the same three genes: APOE, IRS1, and INSR.
One pathway is about insulin signaling breaking down directly. The other is about fat tissue biology going wrong and messing with how APOE gets transported. Two different biological stories but they land on the same molecular players.
That overlap indicates that these three genes sit at a real biological crossroads, where metabolic dysfunction and brain health intersect.
# Approximate median expression values (TPM) read from GTEx violin plots
# Caveat: these are visual estimates for storytelling purposes, not exact downloaded values
tissue_data = {
'Tissue': ['Brain - Hippocampus', 'Brain - Frontal Cortex (BA9)', 'Brain - Amygdala',
'Brain Average (all regions)', 'Adipose - Subcutaneous', 'Liver', 'Muscle - Skeletal'],
'APOE': [1800, 1200, 1000, 1500, 500, 4000, 50],
'IRS1': [5, 4, 4, 5, 10, 5, 10],
'INSR': [25, 20, 18, 20, 30, 10, 40]
}
df_tissue = pd.DataFrame(tissue_data)
df_tissue
| Tissue | APOE | IRS1 | INSR | |
|---|---|---|---|---|
| 0 | Brain - Hippocampus | 1800 | 5 | 25 |
| 1 | Brain - Frontal Cortex (BA9) | 1200 | 4 | 20 |
| 2 | Brain - Amygdala | 1000 | 4 | 18 |
| 3 | Brain Average (all regions) | 1500 | 5 | 20 |
| 4 | Adipose - Subcutaneous | 500 | 10 | 30 |
| 5 | Liver | 4000 | 5 | 10 |
| 6 | Muscle - Skeletal | 50 | 10 | 40 |
# Simple grouped bar chart comparing expression across tissues for the 3 genes
df_plot = df_tissue.set_index('Tissue')
ax = df_plot.plot(kind='bar', figsize=(10, 6), color=['#a83232', '#3268a8', '#32a852'])
plt.title('Expression of Convergent Genes Across Key Tissues')
plt.ylabel('Expression (TPM)')
plt.xlabel('')
plt.xticks(rotation=45, ha='right')
plt.legend(title='Gene')
plt.tight_layout()
plt.savefig('tissue_expression_bars.png', dpi=150, bbox_inches='tight')
plt.show()
The expression data backs this up. APOE is everywhere, it's huge in the liver, but it also shows up a lot in every brain region I looked at, including the hippocampus, which is one of the first areas hit in Alzheimer's. APOE's numbers are so much bigger than IRS1 or INSR that the other two barely show up on the chart, but that's kind of the point, it shows just how much more active APOE is compared to the other two genes. INSR tells a similar story the other way around, it's mostly working in muscle and fat, which makes sense since that's where insulin does its job, but it still shows up in the hippocampus too. So the link isn't just theoretical, you can actually see these genes in both metabolic tissue and brain tissue.
Conclusion¶
Starting from a single question, how does insulin resistance relate to Alzheimer's risk, CFDE REVEAL surfaced two distinct mechanistic hypotheses: one rooted in insulin signaling, the other in adipose lipid transport. Despite different biological logic, both converged on the same three genes: APOE, IRS1, and INSR.
GTEx expression data backs this up. These genes aren't just statistically linked, they're actively expressed in both the metabolic tissues driving insulin resistance and the brain regions most vulnerable in Alzheimer's, including the hippocampus.
This is the kind of thing that's easy to miss when you're searching databases one at a time. Pulling it together in one place made the overlap obvious in a way it probably wouldn't have been otherwise. Granted, the insulin resistance-Alzheimer's link itself isn't new, it's a fairly well-established association at this point, but the point of this exercise was to see if I could rediscover that connection on my own, using CFDE's tools rather than just reading it in a paper. That's really the value of CFDE, it's not generating new data, it's helping you actually see what's already there and trace the evidence for yourself.