<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">

 <title>ergosome</title>
 <link href="https://ergoso.me/atom.xml" rel="self"/>
 <link href="https://ergoso.me"/>
 <updated>2026-01-23T17:33:45+00:00</updated>
 <id>https://ergoso.me</id>
 <author>
   <name>Arman</name>
 </author>

 
 <entry>
   <title>DepMap in the Upside Down: When Tumor Suppressors Reveal Hidden Vulnerabilities</title>
   <link href="https://ergoso.me/depmap/crispr/rbm5/rbm10/activation/lethality/2026/01/23/depmap-upside-down-activation-lethality-rbm5-rbm10.html"/>
   <updated>2026-01-23T01:00:00+00:00</updated>
   <id>https://ergoso.me/depmap/crispr/rbm5/rbm10/activation/lethality/2026/01/23/depmap-upside-down-activation-lethality-rbm5-rbm10</id>
   <content type="html">&lt;p&gt;People (including me) usually come to DepMap for the synthetic lethalities, the strongly selective dependencies, and the extremely negative  gene effect scores (often CERES/Chronos &amp;lt; -1) that scream “this gene is essential for this cell line”; but I have recently found that there is an “Upside Down” view of the data set that is easy to miss.&lt;/p&gt;

&lt;p&gt;This reversed view is anchored based on a handful of “normal-ish” retinal pigment epithelial (RPE) lines and the positive effect observed for some genes that have therapeutic potential to them.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/img/20260123-depmap-upsidedown.jpg&quot; alt=&quot;DepMap - the Upside Down&quot; /&gt;&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;the-usual-direction-negative-gene-effect&quot;&gt;The usual direction: negative gene effect&lt;/h2&gt;

&lt;p&gt;DepMap’s CRISPR gene effect scores are scaled so that a score of 0 corresponds roughly to a typical non-essential gene, while -1 corresponds to the median of common essential genes. More negative values indicate stronger loss of fitness upon knockout. Under the hood, these gene effect values come from pooled CRISPR screens combined with modeling and correction steps. CERES corrects for copy-number driven cutting toxicity, while Chronos models population dynamics to infer relative growth-rate effects over time.&lt;/p&gt;

&lt;p&gt;This framing naturally focuses attention on genes whose loss is deleterious for a good reason (therapeutic target nomination); but, because of this, the positive side of the distribution often stays underexplored.&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;what-does-positive-gene-effect-mean&quot;&gt;What does positive gene effect mean?&lt;/h2&gt;

&lt;p&gt;A positive gene effect means that cells lacking that gene tend to increase in abundance relative to controls over the course of the screen. In other words, the inferred growth rate of the knockout population is higher than the neutral reference.&lt;/p&gt;

&lt;p&gt;There are several non-mutually exclusive reasons this can happen: the most straightforward is true proliferative advantage, where disrupting the gene removes a growth brake. Positive scores can also arise from assay-specific artifacts such as weak guide activity, batch effects, or modeling error. Finally, some effects are small on a per-division basis but accumulate over time in pooled competition assays.&lt;/p&gt;

&lt;p&gt;Importantly, the idea that certain genes act as growth brakes whose loss yields positive selection is not speculative. &lt;a href=&quot;https://www.nature.com/articles/s41467-021-26867-8&quot;&gt;Lenoir et al. previously formalized the concept of proliferation-suppressor genes (PSGs), defined explicitly as genes whose knockout leads to positive selection in CRISPR screens&lt;/a&gt;. They showed that many PSGs are well-known tumor suppressors, including TP53, CDKN1A, CHEK2, and TP53BP1, and that these signals are robust across analysis methods including CERES. Their headline example in AML demonstrates how PSG behavior can uncover pathway-specific liabilities, rather than simply cataloging growth brakes. This supports the idea that positive gene effects can be an organizing signal, not just a curiosity.&lt;/p&gt;

&lt;p&gt;So the intuition is directionally correct. In a sufficiently normal-like context, knocking out a tumor suppressor can look like a growth advantage. The real questions are whether the context is appropriate and whether the signal is consistent enough to trust.&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;enter-the-rpe-lines-normal-ish-with-important-caveats&quot;&gt;Enter the RPE lines: normal-ish with important caveats&lt;/h2&gt;

&lt;p&gt;The RPE lines highlighted here (RPE1SS48, RPE1SS77, RPE1SS6, RPE1SS119, RPE1SS51) are derived from hTERT RPE-1, a widely used non-cancerous retinal pigment epithelial model. According to DepMap and Cellosaurus metadata, these are hTERT-RPE1 derivatives used as part of a controlled panel. RPE1-SS48 and RPE1-SS77 are listed as control clones, while RPE1-SS6, SS51, and SS119 are stable aneuploid derivatives generated by transient treatment with the TTK inhibitor reversine, which induces chromosome missegregation and defined karyotypic changes.&lt;/p&gt;

&lt;p&gt;This distinction matters. These are not primary, unmanipulated normal cells. They are telomerase-immortalized, and some are intentionally aneuploid. As a result, PSG-like signals in these lines may reflect RPE-specific, immortalized, or aneuploid biology rather than a universal normal-cell baseline.&lt;/p&gt;

&lt;p&gt;At the same time, this context is useful precisely because it sits close to the boundary between normal and transformed, while remaining compatible with pooled CRISPR screening.&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;the-upside-down-observation-positive-in-rpe-enriches-for-tumor-suppressors&quot;&gt;The Upside Down observation: positive in RPE enriches for tumor suppressors&lt;/h2&gt;

&lt;p&gt;The observation that &lt;strong&gt;PTEN&lt;/strong&gt; and &lt;strong&gt;APC&lt;/strong&gt; score positively in these RPE lines is consistent with prior PSG work. TP53 shows positive selection in TP53 wild-type backgrounds, and &lt;strong&gt;PTEN&lt;/strong&gt; behaves similarly in &lt;strong&gt;PTEN&lt;/strong&gt; wild-type contexts. A simple framing is that if an RPE subclone is wild-type for a growth brake gene, then knocking it out removes a constraint on proliferation or survival. In pooled assays, that relief can translate into a competitive advantage and a positive gene effect score. This is PSG logic applied to a small set of normal-ish reference lines.&lt;/p&gt;

&lt;p&gt;Taking this idea seriously, we can rank genes by their average positive gene effect across the RPE subclones. When we do this, a familiar pattern emerges. Many of the top-scoring genes are well-known tumor suppressors or checkpoint regulators. Others are less characterized and may represent context-specific growth brakes or indirect effects.&lt;/p&gt;

&lt;p&gt;The table below shows the top 50 such genes, sorted by average RPE score. Genes labeled “Known” are commonly described as tumor suppressors, growth brakes, or checkpoint regulators in at least one major context, often in a context-dependent manner. Genes labeled “Unknown” are not broadly established as tumor suppressors; the accompanying summaries should be read as hypotheses rather than claims.&lt;/p&gt;

&lt;p&gt;Of course, a critical caveat is that these behaviors are observed in the RPE context. Some of these genes may not behave similarly in truly normal primary cells or in other epithelial lineages.&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;gene&lt;/th&gt;
      &lt;th style=&quot;text-align: right&quot;&gt;Average RPE score&lt;/th&gt;
      &lt;th&gt;Relavence to tumor supression&lt;/th&gt;
      &lt;th&gt;Why might depletion be advantageous in RPE (working hypothesis)&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;TP53&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;1.9852569&lt;/td&gt;
      &lt;td&gt;Known&lt;/td&gt;
      &lt;td&gt;Removes a major DNA-damage or apoptosis checkpoint, letting stressed cells keep cycling.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;NF2&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;1.7344470&lt;/td&gt;
      &lt;td&gt;Known&lt;/td&gt;
      &lt;td&gt;Releases contact inhibition and Hippo pathway constraints (YAP and TAZ brake removal).&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;KIRREL1&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;1.2744843&lt;/td&gt;
      &lt;td&gt;Unknown&lt;/td&gt;
      &lt;td&gt;Loss of an adhesion or structure-linked program may reduce contact-dependent growth restraint.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;CDKN1A&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;1.1971946&lt;/td&gt;
      &lt;td&gt;Known&lt;/td&gt;
      &lt;td&gt;Removes p21-mediated cell-cycle arrest downstream of p53 and stress signaling.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;AMOTL2&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.9274090&lt;/td&gt;
      &lt;td&gt;Unknown&lt;/td&gt;
      &lt;td&gt;May weaken Hippo and YAP restraint or junctional growth control in epithelial-like cells.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;PTEN&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.9155450&lt;/td&gt;
      &lt;td&gt;Known&lt;/td&gt;
      &lt;td&gt;Activates PI3K and AKT survival and proliferation signaling by removing a key negative regulator.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;AHR&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.7602595&lt;/td&gt;
      &lt;td&gt;Unknown&lt;/td&gt;
      &lt;td&gt;Disrupts differentiation and xenobiotic-response programs that can restrain proliferation.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;ARNT&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.7455867&lt;/td&gt;
      &lt;td&gt;Unknown&lt;/td&gt;
      &lt;td&gt;Perturbs hypoxia and xenobiotic transcriptional control, which could shift metabolism toward growth.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;USP28&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.6885398&lt;/td&gt;
      &lt;td&gt;Unknown&lt;/td&gt;
      &lt;td&gt;Alters stability of growth regulators and could reduce checkpoint enforcement in certain states.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;DHX29&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.6442140&lt;/td&gt;
      &lt;td&gt;Unknown&lt;/td&gt;
      &lt;td&gt;Translation-initiation stress rewiring could paradoxically favor fast-growing subpopulations.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;AXIN1&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.6126982&lt;/td&gt;
      &lt;td&gt;Known&lt;/td&gt;
      &lt;td&gt;Releases Wnt and beta-catenin control by removing a core negative regulator of Wnt signaling.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;CHEK2&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.6098054&lt;/td&gt;
      &lt;td&gt;Known&lt;/td&gt;
      &lt;td&gt;Weakens DNA damage checkpoints, allowing proliferation despite genomic stress.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;TAOK1&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.6081043&lt;/td&gt;
      &lt;td&gt;Unknown&lt;/td&gt;
      &lt;td&gt;May dampen stress-activated kinase signaling that otherwise slows cell-cycle progression.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;TP53BP1&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.5889426&lt;/td&gt;
      &lt;td&gt;Unknown&lt;/td&gt;
      &lt;td&gt;Reduced DNA repair checkpoint signaling can permit continued cycling after damage.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;RNF146&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.5847380&lt;/td&gt;
      &lt;td&gt;Unknown&lt;/td&gt;
      &lt;td&gt;Could shift Wnt, TNKS, and AXIN turnover dynamics toward proliferative signaling states.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;FBXO42&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.5665235&lt;/td&gt;
      &lt;td&gt;Unknown&lt;/td&gt;
      &lt;td&gt;E3-ligase rewiring may stabilize pro-growth proteins or destabilize growth brakes.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;KEAP1&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.5653373&lt;/td&gt;
      &lt;td&gt;Known&lt;/td&gt;
      &lt;td&gt;Activates NRF2 antioxidant and stress programs by removing NRF2 repression, aiding survival.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;SAV1&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.5534215&lt;/td&gt;
      &lt;td&gt;Known&lt;/td&gt;
      &lt;td&gt;Disables a Hippo pathway scaffold, potentially increasing YAP and TAZ-driven growth programs.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;PTPN14&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.5439258&lt;/td&gt;
      &lt;td&gt;Known&lt;/td&gt;
      &lt;td&gt;Weakens growth restraint via the Hippo and YAP axis and or adhesion-associated signaling.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;LATS2&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.5289316&lt;/td&gt;
      &lt;td&gt;Known&lt;/td&gt;
      &lt;td&gt;Disables a core Hippo kinase, enabling YAP and TAZ pro-growth transcriptional programs.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;BICRA&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.5115757&lt;/td&gt;
      &lt;td&gt;Unknown&lt;/td&gt;
      &lt;td&gt;SWI and SNF remodeling shift could relieve chromatin constraints on proliferation programs.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;IQGAP1&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.5109535&lt;/td&gt;
      &lt;td&gt;Unknown&lt;/td&gt;
      &lt;td&gt;Cytoskeletal and adhesion signaling rewiring may reduce contact-dependent growth restraint.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;KCTD5&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.5074179&lt;/td&gt;
      &lt;td&gt;Unknown&lt;/td&gt;
      &lt;td&gt;Ubiquitin-adaptor effects could alter turnover of growth-inhibitory signaling components.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;FRYL&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.5062773&lt;/td&gt;
      &lt;td&gt;Unknown&lt;/td&gt;
      &lt;td&gt;Loss may alter actin and junction organization, potentially reducing density-dependent arrest.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;NRP1&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.5029067&lt;/td&gt;
      &lt;td&gt;Unknown&lt;/td&gt;
      &lt;td&gt;Growth-factor and co-receptor signaling rewiring may shift pathways toward faster growth states.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;PTPN12&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.4698385&lt;/td&gt;
      &lt;td&gt;Known&lt;/td&gt;
      &lt;td&gt;Removes negative regulation on pro-growth RTK and SRC-family signaling in some contexts.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;ZMAT3&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.4695220&lt;/td&gt;
      &lt;td&gt;Known&lt;/td&gt;
      &lt;td&gt;Disrupts a p53-linked RNA regulation node that can enforce growth restraint.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;PBRM1&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.4627803&lt;/td&gt;
      &lt;td&gt;Known&lt;/td&gt;
      &lt;td&gt;Chromatin remodeling brake removal may relax transcriptional constraints on proliferation.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;CNOT11&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.4569565&lt;/td&gt;
      &lt;td&gt;Unknown&lt;/td&gt;
      &lt;td&gt;mRNA decay and translation control shifts could favor pro-growth transcript programs.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;BRD9&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.4438428&lt;/td&gt;
      &lt;td&gt;Unknown&lt;/td&gt;
      &lt;td&gt;Non-canonical BAF remodeling changes could relax differentiation constraints in RPE state.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;R3HDM4&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.4420323&lt;/td&gt;
      &lt;td&gt;Unknown&lt;/td&gt;
      &lt;td&gt;RNA metabolism changes could tilt expression toward proliferative isoforms and programs.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;CCDC159&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.4351946&lt;/td&gt;
      &lt;td&gt;Unknown&lt;/td&gt;
      &lt;td&gt;Poorly characterized; could reflect lineage-specific growth restraint circuitry.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;PDCD10&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.4319345&lt;/td&gt;
      &lt;td&gt;Unknown&lt;/td&gt;
      &lt;td&gt;Perturbs apoptosis and vascular signaling modules; may reduce stress-induced death.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;FAM193A&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.4306706&lt;/td&gt;
      &lt;td&gt;Unknown&lt;/td&gt;
      &lt;td&gt;Poorly characterized; could mark an RPE-specific brake or assay-correlated effect.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;APC&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.4280347&lt;/td&gt;
      &lt;td&gt;Known&lt;/td&gt;
      &lt;td&gt;Releases Wnt and beta-catenin signaling control by removing a core tumor suppressor brake.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;BAG6&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.4258618&lt;/td&gt;
      &lt;td&gt;Unknown&lt;/td&gt;
      &lt;td&gt;Alters proteostasis, immune presentation, and apoptosis balance, potentially improving survival.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;HLA-DQB1&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.4235409&lt;/td&gt;
      &lt;td&gt;Unknown&lt;/td&gt;
      &lt;td&gt;Likely not a true growth brake; could reflect expression and fitness model edge cases in vitro.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;TRAF3&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.4225509&lt;/td&gt;
      &lt;td&gt;Known&lt;/td&gt;
      &lt;td&gt;Loss can activate NF-kB survival signaling in certain cellular contexts.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;CCDC6&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.4087856&lt;/td&gt;
      &lt;td&gt;Known&lt;/td&gt;
      &lt;td&gt;Disrupts DNA damage response and growth restraint functions reported in multiple cancers.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;CHD8&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.3933085&lt;/td&gt;
      &lt;td&gt;Unknown&lt;/td&gt;
      &lt;td&gt;Chromatin regulation shift may de-repress pro-proliferative transcriptional programs.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;BABAM1&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.3927088&lt;/td&gt;
      &lt;td&gt;Unknown&lt;/td&gt;
      &lt;td&gt;DNA repair complex perturbation could weaken checkpoints that slow cycling after damage.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;PPM1F&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.3913650&lt;/td&gt;
      &lt;td&gt;Unknown&lt;/td&gt;
      &lt;td&gt;Phosphatase loss may sustain kinase-driven pro-growth signaling states.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;RB1&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.3900008&lt;/td&gt;
      &lt;td&gt;Known&lt;/td&gt;
      &lt;td&gt;Removes G1 and S restriction point control, enabling unchecked cell-cycle progression.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;ZNF853&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.3780378&lt;/td&gt;
      &lt;td&gt;Unknown&lt;/td&gt;
      &lt;td&gt;Transcriptional regulation shift; could remove lineage-specific growth restraint.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;ATXN3L&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.3769404&lt;/td&gt;
      &lt;td&gt;Unknown&lt;/td&gt;
      &lt;td&gt;Deubiquitinase-like effects may alter stability of growth-regulatory proteins.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;CAND1&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.3744396&lt;/td&gt;
      &lt;td&gt;Unknown&lt;/td&gt;
      &lt;td&gt;Cullin-RING ligase cycling changes could stabilize proteins that promote proliferation.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;ARHGAP35&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.3673934&lt;/td&gt;
      &lt;td&gt;Unknown&lt;/td&gt;
      &lt;td&gt;Rho GTPase signaling rewiring may alter adhesion and contact inhibition toward faster growth.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;CASD1&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.3631760&lt;/td&gt;
      &lt;td&gt;Unknown&lt;/td&gt;
      &lt;td&gt;Lipid and glycosylation changes might alter membrane signaling to favor growth.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;NAPB&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.3608647&lt;/td&gt;
      &lt;td&gt;Unknown&lt;/td&gt;
      &lt;td&gt;Vesicle trafficking changes could shift receptor recycling and signaling toward proliferative states.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;KDM6A&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.3554378&lt;/td&gt;
      &lt;td&gt;Known&lt;/td&gt;
      &lt;td&gt;Epigenetic brake removal (H3K27 demethylase context) can favor proliferation in some settings.&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;from-psgs-in-rpe-to-activation-lethality-in-cancer&quot;&gt;From PSGs in RPE to activation lethality in cancer&lt;/h2&gt;

&lt;p&gt;The conceptual extension is straightforward. If loss of a gene confers a proliferative advantage in a normal-like context, then cancers that effectively live downstream of that loss may rely on compensatory pathways to survive the resulting rewired state. Those compensatory mechanisms can become druggable dependencies.&lt;/p&gt;

&lt;p&gt;This is the mirror image of synthetic lethality. Instead of two losses killing the cell, removal of a brake pushes the system into a stressed configuration that requires new support.&lt;/p&gt;

&lt;p&gt;Operationally, this suggests a filter strategy that starts with genes that are positive in RPE, removes obvious confounders such as immune-lineage artifacts or low-quality screen signals, requires near-neutral behavior across solid tumor lines on average, and then focuses on genes that show strong dependency in a subset of cancer lines. In other words, the target is not the proliferation suppressor itself, but the load-bearing dependencies that emerge downstream. Below is a table that lists genes that meet these criteria:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Gene&lt;/th&gt;
      &lt;th style=&quot;text-align: right&quot;&gt;Average RPE score&lt;/th&gt;
      &lt;th style=&quot;text-align: right&quot;&gt;Average score across solid lines&lt;/th&gt;
      &lt;th style=&quot;text-align: right&quot;&gt;Number of dependent lines&lt;/th&gt;
      &lt;th&gt;Heme or myeloid bias&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;FBXO42&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.57&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;-0.16&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;1&lt;/td&gt;
      &lt;td&gt;No&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;KEAP1&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.57&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;-0.15&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;13&lt;/td&gt;
      &lt;td&gt;No&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;TSC2&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.33&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;-0.05&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;1&lt;/td&gt;
      &lt;td&gt;Yes&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;RBM5&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.33&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;-0.20&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;4&lt;/td&gt;
      &lt;td&gt;No&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;CREBBP&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.32&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;-0.11&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;10&lt;/td&gt;
      &lt;td&gt;No&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;AGPAT5&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.28&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.08&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;1&lt;/td&gt;
      &lt;td&gt;No&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;SMARCA4&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.28&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;-0.35&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;3&lt;/td&gt;
      &lt;td&gt;No&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;FLI1&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.25&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;-0.04&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;8&lt;/td&gt;
      &lt;td&gt;Yes&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;EP300&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.25&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;-0.26&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;6&lt;/td&gt;
      &lt;td&gt;Yes&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;CAB39&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.25&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;-0.18&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;2&lt;/td&gt;
      &lt;td&gt;Yes&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;CBFB&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.23&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;-0.15&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;5&lt;/td&gt;
      &lt;td&gt;Yes&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;LARS2&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.21&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;-0.48&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;11&lt;/td&gt;
      &lt;td&gt;No&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;MYBL1&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.20&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;0.10&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;1&lt;/td&gt;
      &lt;td&gt;No&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Among these, there are a few familiar faces – for example, the &lt;strong&gt;CREBBP&lt;/strong&gt; and &lt;strong&gt;EP300&lt;/strong&gt; activator pair:&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/img/20260123-depmap-crebbp-ep300.jpg&quot; alt=&quot;CREBBP and EP300 activation lethalithy&quot; /&gt;&lt;/p&gt;

&lt;p&gt;and some trivial ones – for example, &lt;strong&gt;FLI1&lt;/strong&gt; overexpression creating a vulnerability for &lt;strong&gt;FLI1&lt;/strong&gt; depletion:&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/img/20260123-depmap-fli1.jpg&quot; alt=&quot;FLI1 overexpression leads to FLI1 dependency&quot; /&gt;&lt;/p&gt;

&lt;p&gt;When you exlude the lymph- or myeloid-lineage specific vulnerabilities, the list gets even shorter and it gets easier to explore the profiles manually to see if there are interesting stories behind those. While doing that, &lt;strong&gt;RBM5&lt;/strong&gt; caught my eye because it has a very interesting DepMap profile (i.e. enrichment of a specific lineage and a related feature that strongly correlates with the dependency):&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/img/20260123-depmap-rbm5-profile.jpg&quot; alt=&quot;RBM10-loss leads to RBM5 dependency&quot; /&gt;&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;power-play-between-rbm5-and-rbm10&quot;&gt;Power play between RBM5 and RBM10&lt;/h2&gt;

&lt;p&gt;Based on the DepMap profile, in RBM10-loss cell lines, RBM5 appears to function as a load-bearing compensator. At first glance, this might suggest a straightforward synthetic lethality, given that both genes encode RNA-binding proteins and share partially overlapping regulatory space. However, the data and the broader context argue for something slightly different.&lt;/p&gt;

&lt;p&gt;Rather than behaving like a classic synthetic lethal partner, RBM5 fits the pattern of a compensatory node that becomes essential only after the system has been pushed into a rewired state. In normal-like RPE cells, loss of RBM5 is tolerated and even associated with a proliferative advantage. In contrast, in RBM10-deficient cancer cells, RBM5 appears to sit at a point of fragility. If this interpretation is correct, disrupting RBM5 in an RBM10-loss context does not simply remove redundancy, but instead pushes an already stressed regulatory system past its tolerance threshold, producing a kill signal consistent with activation lethality.&lt;/p&gt;

&lt;p&gt;At this point, caution is warranted. Many vulnerabilities identified in cell lines turn out to be rare or clinically irrelevant once patient genomics are considered. This is where the RBM5 story becomes more interesting. RBM10 loss is not a rare event. It is recurrent across multiple tumor types, including non-small cell lung cancer, colorectal cancer, bladder cancer, and renal cancer. Moreover, RBM10 loss has been shown to promote EGFR-driven lung cancer, indicating that this alteration is selected for and functionally relevant rather than a neutral passenger event.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/img/20260123-depmap-rbm5.jpg&quot; alt=&quot;RBM10-loss leads to RBM5 dependency&quot; /&gt;&lt;/p&gt;

&lt;p&gt;That context matters. It suggests that the genomic background required for this potential activation lethality is present in real patient populations.&lt;/p&gt;

&lt;p&gt;At the same time, directly targeting a spliceosomal regulator is inherently risky. Disrupting RBM5 globally would likely perturb core RNA-processing programs in normal cells, and as we know, RBM5 loss is associated with increased proliferation in RPE cells. However, this does not mean the biology is untargetable.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://pmc.ncbi.nlm.nih.gov/articles/PMC5127646/&quot;&gt;RBM5 is known to tune alternative splicing of apoptosis and stress-response transcripts, including FAS/CD95 and CASP2, by interfering with late splice-site pairing&lt;/a&gt;. RBM5 also counterbalances RBM10 activity at key exons such as NUMB exon 9, a regulator of NOTCH signaling. In turn, &lt;a href=&quot;https://academic.oup.com/nar/article/45/14/8524/3861608&quot;&gt;RBM10 down-regulates RBM5 through alternative splicing coupled to nonsense-mediated decay&lt;/a&gt;, linking their dosages. Together, these observations suggest that RBM10 loss reshapes a specific splicing landscape rather than creating a generic dependence on splicing itself.&lt;/p&gt;

&lt;p&gt;This raises the possibility that therapeutic intervention may not require direct inhibition of RBM5. Instead, it may be possible to identify downstream transcripts or splice isoforms that are aberrantly regulated and uniquely required in RBM10-loss contexts, and to target those vulnerabilities using alternative modalities.&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;closing-looking-at-depmap-from-the-other-side&quot;&gt;Closing: Looking at DepMap from the other side&lt;/h2&gt;

&lt;p&gt;Most analyses of functional genomics data are optimized to find the obvious signals: strong dependencies, clean synthetic lethalities, and genes whose loss is catastrophic. That approach has been extraordinarily productive, and it should remain the default.&lt;/p&gt;

&lt;p&gt;But datasets like DepMap are rich enough to reward curiosity in other directions. Sometimes, by looking at what helps cells grow rather than what hurts them, we can surface growth brakes, compensatory states, and fragile rewired systems that would otherwise remain invisible. There are not many such leads, and most will not survive scrutiny. But a few do, and they often point to biology that sits just outside the standard playbook.&lt;/p&gt;

&lt;p&gt;The Upside Down of DepMap is not a replacement for conventional analysis. It is a reminder that complex systems rarely advertise their weaknesses directly. Sometimes, the most interesting vulnerabilities are hiding in plain sight, waiting for someone to look at the data from an unusual angle.&lt;/p&gt;
</content>
 </entry>
 
 <entry>
   <title>FAS: The Death Receptor That Keeps the Immune System in Balance</title>
   <link href="https://ergoso.me/ai/vibe/research/fas/fasl/cd95/apo1/tcr/bcr/2026/01/05/gene-stories-tcr-fas-cd95.html"/>
   <updated>2026-01-05T00:00:01+00:00</updated>
   <id>https://ergoso.me/ai/vibe/research/fas/fasl/cd95/apo1/tcr/bcr/2026/01/05/gene-stories-tcr-fas-cd95</id>
   <content type="html">&lt;table&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;a href=&quot;/ai/vibe/research/rag1/rag2/vdj/tcr/bcr/2025/12/29/gene-stories-tcr-rag-rag1-rag2.html&quot;&gt;Part I&lt;/a&gt;&lt;/td&gt;
      &lt;td&gt;&lt;a href=&quot;/ai/vibe/research/aire/rag2/vdj/tcr/bcr/2026/01/01/gene-stories-tcr-aire.html&quot;&gt;Part II&lt;/a&gt;&lt;/td&gt;
      &lt;td&gt;Part III (you are reading this)&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;If you have ever watched a crowd surge toward an emergency, sirens blaring and lights flashing, you have seen the good side of a fast response. But the real test of any system is not only how quickly it mobilizes. It is whether it can stand down once the crisis is over.&lt;/p&gt;

&lt;p&gt;The immune system faces this exact challenge: it can expand armies of lymphocytes in days, but when the threat passes, those armies must shrink cleanly and quietly. &lt;strong&gt;FAS&lt;/strong&gt; (also known as &lt;strong&gt;CD95&lt;/strong&gt; or &lt;strong&gt;APO-1&lt;/strong&gt;) is one of the signals that makes this possible. When it works, immune responses resolve and self-tolerance holds. When it fails, immune cells linger like responders who never go home, filling lymph nodes, spilling into tissues, and sometimes turning their weapons inward.&lt;/p&gt;

&lt;p&gt;This post focuses on peripheral tolerance, where the immune system actively deletes activated or autoreactive cells that escaped earlier checkpoints. Few pathways illustrate the necessity of timely disposal better than Fas, an accidental antibody finding that became a cornerstone of immunology, mouse genetics, and human disease.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/img/20260105-timeline.jpg&quot; alt=&quot;FAS story timeline&quot; /&gt;&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;two-antibodies-two-continents-one-unsettling-observation-1989&quot;&gt;Two antibodies, two continents, one unsettling observation (1989)&lt;/h2&gt;

&lt;p&gt;In the late 1980s, monoclonal antibodies were already transforming biology. Most were used as labels or blockers: bind a protein, stop a ligand, map a marker. Then something unexpected happened in Tokyo.&lt;/p&gt;

&lt;p&gt;Shin Yonehara and colleagues generated a monoclonal antibody, often referred to as CH-11, that did not merely bind a cell-surface protein. It killed the cell. Even more striking, the killing was enhanced when cells were sensitized with actinomycin D, a clue that what they were triggering was not nonspecific damage but a regulated program the cell could execute when survival transcripts were suppressed. In 1989, Yonehara reported this phenomenon and named the antigen “Fas” after the FS-7 cell line used during antibody generation, an oddly mundane origin for a receptor that would become synonymous with cellular suicide.&lt;/p&gt;

&lt;p&gt;Almost simultaneously in Germany, Peter H. Krammer’s group, with Bernhard Trauth as first author, described an agonistic antibody against a surface molecule they called APO-1. Their antibody also triggered programmed cell death in tumor cells. The convergence was uncanny: two independent laboratories on different continents had found that binding a particular surface receptor could function as an off switch for cell survival.&lt;/p&gt;

&lt;p&gt;At the time, discovering such antibodies was anything but straightforward. Researchers immunized mice with whole human cells, fused thousands of antibody-producing B cells into hybridomas, and then screened tens of thousands of clones by hand. Crucially, these screens were functional rather than descriptive: instead of asking whether an antibody bound a cell, they asked whether it did something dramatic, like halting growth or killing the cell outright. With no high-throughput assays, no genomics, and no rapid way to identify antibody targets, recognizing that a rare antibody was triggering a regulated death program rather than nonspecific toxicity required painstaking controls, intuition, and persistence. Confirming that the death was apoptosis relied on morphology, membrane blebbing, DNA fragmentation, and the absence of inflammation. Apoptosis is a tidy self-destruct mechanism. DNA is chopped into fragments, the membrane forms blebs, and the cell breaks into apoptotic bodies that can be cleared without triggering inflammatory chaos. In retrospect, the discovery of Fas and APO-1 looks inevitable, but at the time it was a low-probability find that depended as much on experimental courage as on technical skill.&lt;/p&gt;

&lt;p&gt;By the end of 1989, the field was hooked. If Fas or APO-1 was a genuine death-inducing receptor, the questions were obvious: what was this receptor, why did immune cells carry it, and whatphysiological conditions triggered its expression?&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;from-antigen-to-gene-and-the-lpr-mouse-mystery-1991-1992&quot;&gt;From antigen to gene and the lpr mouse mystery (1991-1992)&lt;/h2&gt;

&lt;h3 id=&quot;cloning-fas-in-the-pre-genomic-era-1991&quot;&gt;Cloning Fas in the pre-genomic era (1991)&lt;/h3&gt;

&lt;p&gt;The next milestone required molecular cloning, and it reflected both the limits and ingenuity of early 1990s biology. In 1988, Yonehara approached Shigekazu Nagata, already known for work on cytokine receptors, to help isolate the gene encoding Fas. At the time, cloning meant building and screening cDNA libraries and validating candidates experimentally, not simply sequencing a sample and searching a database. The team had to build cDNA libraries from Fas-positive cells, express them in mammalian cells, and repeatedly enrich for the rare transfectants that gained anti-Fas antibody staining by flow cytometry. Early library screens even failed, forcing a switch to stronger mammalian expression systems.&lt;/p&gt;

&lt;p&gt;In 1991, Naoto Itoh, Yonehara, Nagata, and colleagues finally cloned the human Fas cDNA and showed that Fas belongs to the TNF receptor superfamily. The protein carried a cysteine-rich extracellular domain and a cytoplasmic death domain. When expressed in cultured cells, Fas could mediate apoptosis: cells transfected with Fas would self-destruct when exposed to agonistic anti-Fas antibody. This turned a strange antibody effect into a new category of signaling molecule, the death receptor.&lt;/p&gt;

&lt;h3 id=&quot;lpr-mice-were-not-hyperactive-they-were-failing-to-die-1992&quot;&gt;lpr mice were not hyperactive, they were failing to die (1992)&lt;/h3&gt;

&lt;p&gt;At the same time, immunologists had long been puzzled by a spontaneous mutant mouse strain called &lt;strong&gt;lpr&lt;/strong&gt;, short for lymphoproliferation. Especially on autoimmune-prone backgrounds such as MRL, lpr mice developed massive lymph node and spleen enlargement, produced autoantibodies reminiscent of lupus, and suffered immune-complex kidney disease. Their lymph nodes filled with abnormal CD4-negative, CD8-negative double-negative T cells, an unusual population that became one of the most recognizable clues in immune tolerance genetics.&lt;/p&gt;

&lt;p&gt;Mapping a mouse mutation in that era was slow and laborious. There was no whole-genome sequencing and no easy way to recreate candidate mutations to test causality. Geneticists had narrowed lpr to mouse chromosome 19, but the responsible gene remained unknown. Nagata’s newly cloned Fas gene offered a compelling candidate. In 1992, Rie Watanabe-Fukunaga in Nagata’s lab, working with NIH mouse geneticists Nancy Jenkins and Neal Copeland, examined Fas expression in lpr mice. The result was decisive: lpr mice showed almost no functional Fas expression, and a genetic lesion in the Fas gene explained the phenotype. Later work revealed different lpr alleles, including a retrotransposon insertion that disrupted transcription and a point mutation in the death domain, known as the lpr^cg allele, that rendered Fas nonfunctional.&lt;/p&gt;

&lt;p&gt;This finding forged a direct link between defective apoptosis and autoimmunity. If activated T cells cannot receive the Fas time-to-die signal, they persist longer than they should, accumulate in lymphoid organs, and increase the likelihood of autoreactivity and tissue damage. The 1992 Nature paper went further, noting that Fas is normally expressed in the thymus and likely plays an important role in deleting autoreactive T cells, bridging central and peripheral tolerance.&lt;/p&gt;

&lt;h3 id=&quot;how-fas-kills-death-by-design&quot;&gt;How Fas kills: death by design&lt;/h3&gt;

&lt;p&gt;Mechanistically, Fas signaling turned out to be both simple and ruthless. When Fas is engaged by ligand or agonistic antibody, it trimerizes and assembles the death-inducing signaling complex, or DISC. The cytoplasmic death domain recruits the adaptor protein FADD, which in turn recruits procaspase-8, historically also called FLICE. Procaspase-8 molecules activate one another by proximity, then trigger downstream executioner caspases such as caspase-3. DNA fragments, cellular scaffolding collapses, and the cell is dismantled from within, usually without inflammation.&lt;/p&gt;

&lt;p&gt;By the early 1990s, the picture was coming together: a TNF-family receptor on lymphocytes that could trigger apoptosis, a mutant mouse lacking this receptor that developed autoimmune lymphoproliferation, and a mechanistic pathway linking receptor engagement to caspase activation. One crucial piece was still missing: the physiological ligand.&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;fas-ligand-and-the-gld-mirror-image-1993-1994&quot;&gt;Fas ligand and the gld mirror image (1993-1994)&lt;/h2&gt;

&lt;h3 id=&quot;the-ligand-appears-1993&quot;&gt;The ligand appears (1993)&lt;/h3&gt;

&lt;p&gt;Every receptor needs a ligand. For Fas, the search concluded in late 1993 when Nagata’s group identified and cloned Fas Ligand, another TNF-family molecule expressed on activated T cells and natural killer cells. This discovery made Fas biologically intuitive. Cytotoxic lymphocytes could eliminate targets by engaging Fas, and immune responses could terminate through activation-induced cell death, a contraction phase that prevents chronic immune activation and preserves self-tolerance.&lt;/p&gt;

&lt;h3 id=&quot;gld-mice-complete-the-genetic-picture-1994&quot;&gt;gld mice complete the genetic picture (1994)&lt;/h3&gt;

&lt;p&gt;At nearly the same time, attention returned to another autoimmune mouse strain: gld, for generalized lymphoproliferative disease. gld mice developed a phenotype almost indistinguishable from lpr mice, yet the mutation mapped elsewhere. In 1994, researchers showed that gld mice carry a point mutation in the Fas ligand gene, producing a nonfunctional ligand unable to trigger Fas-mediated death.&lt;/p&gt;

&lt;p&gt;The symmetry was striking. lpr mice lacked a functional Fas receptor. gld mice lacked a functional Fas ligand. Different mutations, same outcome: immune cells that should die persisted, lymphoid organs enlarged, and autoimmunity followed. Together, these mutants established the Fas and FasL pair as indispensable for immune regulation and tolerance.&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;a-human-syndrome-finds-its-molecular-cause-1967-1995&quot;&gt;A human syndrome finds its molecular cause (1967-1995)&lt;/h2&gt;

&lt;p&gt;Long before Fas was known, clinicians had described a rare childhood disorder characterized by chronic, non-malignant lymphadenopathy, splenomegaly, and autoimmune cytopenias, without infection or cancer. Frank Canale and Robert Smith first reported it in 1967. For decades, it remained a medical curiosity.&lt;/p&gt;

&lt;p&gt;By the mid-1990s, immunologists had the tools and conceptual framework to connect this syndrome to the Fas pathway. In 1995, Frédéric Rieux-Laucat and Alain Fischer in Paris identified mutations in the FAS gene in children with Canale-Smith syndrome. That same year, an NIH team led by Jennifer Puck, Steven Straus, and Michael Lenardo independently identified FAS mutations in American patients with similar clinical features. The disorder was renamed Autoimmune Lymphoproliferative Syndrome, or ALPS. A key genetic insight was that many cases involve dominant-negative FAS mutations. A single mutated allele can sabotage Fas signaling even when the other allele is normal, explaining autosomal dominant inheritance in many families.&lt;/p&gt;

&lt;p&gt;Clinically, ALPS often presents early in childhood with persistent enlarged lymph nodes and spleen, along with autoimmune destruction of blood cells such as hemolytic anemia and thrombocytopenia. A striking laboratory hallmark is the expansion of CD3-positive double-negative T cells, sometimes comprising 15 to 30 percent of circulating T cells. These cells mirror the abnormal population seen in lpr mice and have become a diagnostic signature of the disease.&lt;/p&gt;

&lt;p&gt;ALPS also highlighted a deep connection between apoptosis and cancer risk. When lymphocytes that should die persist, they not only drive autoimmunity but also increase the chance of malignant transformation. Patients with ALPS have an elevated risk of lymphoma. The connection works in the opposite direction as well: many tumors evade immune attack by downregulating Fas or exploiting FasL to eliminate infiltrating T cells.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/img/20260105-overview.jpg&quot; alt=&quot;Overview of the FAS story&quot; /&gt;&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;epilogue-from-accidental-discovery-to-modern-immunology-1996-present&quot;&gt;Epilogue: From accidental discovery to modern immunology (1996-Present)&lt;/h2&gt;

&lt;p&gt;The Fas story is a textbook example of biology progressing from accident to mechanism to medicine. In 1989, antibody-induced cell death looked like a laboratory oddity. In the early 1990s, Fas and FasL became molecular entities with defined domains, defined signaling partners, and defined genetics, validated by two classic mouse mutants. By 1995, the same pathway explained a rare human disease, transforming a clinical curiosity into a diagnosable genetic syndrome.&lt;/p&gt;

&lt;p&gt;Since then, the picture has only grown more nuanced: we now know that Fas signaling is not always a one-way route to apoptosis. Depending on cellular context and downstream pathway integrity, CD95 can also engage non-apoptotic signaling, including inflammatory or survival pathways. This complexity helps explain why directly activating Fas has proven risky as a therapeutic strategy, since some tissues, particularly the liver, are exquisitely sensitive to death-receptor signaling.&lt;/p&gt;

&lt;p&gt;Immune tolerance is not a single mechanism but a layered system. When one layer fails, the result can be immunological chaos. Imagine a self-reactive clonotype that somehow slips through thymic negative selection, which happens from time to time because AIRE dramatically expands the self-antigen “training set,” but the thymus is not an omniscient filter, and some potentially problematic T cells still graduate from thymus. That is where Fas starts to matter most: once a T cell is repeatedly stimulated, especially in a chronic setting like persistent self-antigen exposure (leading to autoimmunity), the immune system can invoke a built-in off switch called activation-induced cell death. In effect, Fas helps ensure that a self-reactive T cell that “passed the exam” but misbehaves later does not get to stick around indefinitely.&lt;/p&gt;

&lt;p&gt;So, in the case of self-reactive T cells and an immune system that cannot stop responding, Fas has taught immunology one of its most counterintuitive truths: controlled cell death is essential for a proper immune system.&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;references&quot;&gt;References&lt;/h2&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.pharm.kyoto-u.ac.jp/Yonehara_Lab/SY_Nomenclature.htm&quot;&gt;Yonehara S. et al., 1989 – First identification of Fas antigen via a cell-killing monoclonal antibody&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://europepmc.org/article/MED/2787530&quot;&gt;Trauth B. et al., 1989 – Monoclonal antibody-mediated tumor regression by induction of apoptosis (APO-1) (&lt;em&gt;Science&lt;/em&gt;)&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.cell.com/cell/abstract/0092-8674%2891%2990614-5&quot;&gt;Itoh N., Nagata S. et al., 1991 – The polypeptide encoded by the cDNA for Fas can mediate apoptosis (&lt;em&gt;Cell&lt;/em&gt;)&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://pubmed.ncbi.nlm.nih.gov/1372394/&quot;&gt;Watanabe-Fukunaga R. et al., 1992 – Lymphoproliferation disorder in mice explained by defects in Fas antigen (&lt;em&gt;Nature&lt;/em&gt;)&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://pubmed.ncbi.nlm.nih.gov/7511063/&quot;&gt;Takahashi T., Nagata S. et al., 1994 – Generalized lymphoproliferative disease caused by Fas ligand point mutation (&lt;em&gt;gld&lt;/em&gt;)&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://pubmed.ncbi.nlm.nih.gov/7539157/&quot;&gt;Rieux-Laucat F., Fischer A. et al., 1995 – Mutations in Fas associated with human lymphoproliferative syndrome and autoimmunity (&lt;em&gt;Science&lt;/em&gt;)&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.researchgate.net/profile/Madhu-Ramaswamy/publication/233738787_Autoimmunity_Twenty_Years_in_the_Fas_Lane/links/569408c708ae425c689624f4/Autoimmunity-Twenty-Years-in-the-Fas-Lane.pdf&quot;&gt;Ramaswamy M. &amp;amp; Siegel R., 2012 – “Twenty Years in the Fas Lane” review (&lt;em&gt;J. Immunol.&lt;/em&gt;)&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.cytometry.org/newsletters/eICCS-2-1/article5.php&quot;&gt;ICCS Newsletter, 2011 – Overview of ALPS clinical features and Fas pathway defects&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://atlasgeneticsoncology.org/deep-insight/20040/the-fas-fas-ligand-apoptotic-pathway&quot;&gt;Atlas of Genetics (Pierre Bobé, 2002) – The Fas–FasL apoptotic pathway and DISC formation&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.frontiersin.org/journals/cellular-and-infection-microbiology/articles/10.3389/fcimb.2025.1561102/full&quot;&gt;Hu et al., 2025 – Fas mediates apoptosis, inflammation, and host defense&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.accessdata.fda.gov/drugsatfda_docs/label/2022/021083s069s070%2C021110s087s088lbl.pdf&quot;&gt;FDA label – Rapamune (sirolimus)&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.niaid.nih.gov/diseases-conditions/autoimmune-lymphoproliferative-syndrome-treatment&quot;&gt;NIAID – Autoimmune Lymphoproliferative Syndrome (ALPS) Treatment&lt;/a&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;hr /&gt;

&lt;table&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;a href=&quot;/ai/vibe/research/rag1/rag2/vdj/tcr/bcr/2025/12/29/gene-stories-tcr-rag-rag1-rag2.html&quot;&gt;Part I&lt;/a&gt;&lt;/td&gt;
      &lt;td&gt;&lt;a href=&quot;/ai/vibe/research/aire/rag2/vdj/tcr/bcr/2026/01/01/gene-stories-tcr-aire.html&quot;&gt;Part II&lt;/a&gt;&lt;/td&gt;
      &lt;td&gt;Part III (you are reading this)&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;
</content>
 </entry>
 
 <entry>
   <title>AIRE: The Orchestrator of Self-Tolerance</title>
   <link href="https://ergoso.me/ai/vibe/research/aire/rag2/vdj/tcr/bcr/2026/01/01/gene-stories-tcr-aire.html"/>
   <updated>2026-01-01T00:00:01+00:00</updated>
   <id>https://ergoso.me/ai/vibe/research/aire/rag2/vdj/tcr/bcr/2026/01/01/gene-stories-tcr-aire</id>
   <content type="html">&lt;table&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;a href=&quot;/ai/vibe/research/rag1/rag2/vdj/tcr/bcr/2025/12/29/gene-stories-tcr-rag-rag1-rag2.html&quot;&gt;Part I&lt;/a&gt;&lt;/td&gt;
      &lt;td&gt;Part II (you are reading this)&lt;/td&gt;
      &lt;td&gt;&lt;a href=&quot;/ai/vibe/research/fas/fasl/cd95/apo1/tcr/bcr/2026/01/05/gene-stories-tcr-fas-cd95.html&quot;&gt;Part III&lt;/a&gt;&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Your immune system is a powerful guard dog, but it needs training. From the moment you are born, it has to learn the difference between intruders and you. Most of us never notice that training happening. But in rare cases, the lesson plan goes unexpectedly wrong, and the immune system grows up confused, lashing out at the body while letting some infections slip through.&lt;/p&gt;

&lt;p&gt;This post follows the trail from baffling patient cases to an unexpected culprit called &lt;strong&gt;AIRE&lt;/strong&gt;, a tiny genetic teacher that helps the immune system recognize itself.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/img/20260101-timeline.jpg&quot; alt=&quot;AIE story timeline&quot; /&gt;&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;the-mystery-of-the-missing-immune-tolerance-1926-to-1946&quot;&gt;The Mystery of the Missing Immune Tolerance (1926 to 1946)&lt;/h2&gt;

&lt;p&gt;In the early 20th century, doctors began noticing a strange pattern in a small number of children. These patients did not just have one problem. They had a cluster of them: thrush infections that kept coming back no matter what clinicians tried; dangerously low calcium from failing parathyroid glands, sometimes severe enough to trigger seizures; and later, signs of Addison’s disease, when the adrenal glands stopped producing essential hormones, often accompanied by a deep bronzed pigmentation that looked almost like an unnatural tan.&lt;/p&gt;

&lt;p&gt;Individually, each diagnosis was recognizable. Together, they were baffling. In 1926, Dr. Ernest Schmidt in Germany noted the curious association between persistent Candida infections and underactive parathyroid glands. By 1946, clinicians recognized that these issues often appeared alongside autoimmune adrenal failure, defining a syndrome that seemed to combine immune weakness and immune misdirection in the same person.&lt;/p&gt;

&lt;p&gt;Much later, this condition would be given two names that referred to the same disease: &lt;strong&gt;APS-1 (Autoimmune Polyglandular Syndrome Type 1)&lt;/strong&gt; and &lt;strong&gt;APECED (Autoimmune Polyendocrinopathy-Candidiasis-Ectodermal Dystrophy)&lt;/strong&gt;. The names captured the breadth of what patients experienced, but they did not solve the underlying question: why would the immune system attack multiple organs while still failing to control a common fungus?&lt;/p&gt;

&lt;h3 id=&quot;how-the-syndrome-came-into-focus-1980s-to-1990&quot;&gt;How the Syndrome Came Into Focus (1980s to 1990)&lt;/h3&gt;

&lt;p&gt;As more cases accumulated, the outline of the disorder became clearer, especially through the work of Finnish pediatric endocrinologist Jaakko Perheentupa. In an era when rare diseases were often described through isolated case reports, Perheentupa did something unusually patient: he followed families over years, sometimes decades, watching the syndrome unfold rather than treating it as a static diagnosis.&lt;/p&gt;

&lt;p&gt;In a landmark 1990 study, he and colleagues documented 68 cases, many in Finland, helping establish what APS-1 looked like over time rather than at a single moment. One insight stood out immediately: the syndrome rarely arrived all at once. Many patients developed two major features in childhood, with the third appearing later. This staging explained why earlier reports had seemed inconsistent or incomplete.&lt;/p&gt;

&lt;p&gt;The second insight was genetic. APS-1 behaved like an autosomal recessive condition, clustering in families and appearing more often in communities with shared ancestry. Clinicians converged on a practical diagnostic rule: at least two of the classic triad, candidiasis, hypoparathyroidism, and Addison’s disease.&lt;/p&gt;

&lt;p&gt;By the early 1980s, researchers also recognized that this syndrome was distinct from a more common adult-onset disorder that also included Addison’s disease. In 1981, immunologist Markus Neufeld formalized the naming: APS Type 1 for the childhood-onset triad, and APS Type 2 for adult-onset Addison’s plus other autoimmune endocrinopathies. Around the same period, the term APECED gained popularity because it highlighted features beyond the endocrine glands, especially ectodermal changes affecting teeth, nails, skin, and hair. Other labels appeared briefly, but APS-1 and APECED became the enduring names.&lt;/p&gt;

&lt;h3 id=&quot;aps-1-vs-apeced-whats-in-a-name&quot;&gt;APS-1 vs APECED: What’s in a Name?&lt;/h3&gt;

&lt;p&gt;APS-1 and APECED describe the same disease.&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Term&lt;/th&gt;
      &lt;th&gt;Emphasis&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;APS-1&lt;/td&gt;
      &lt;td&gt;Autoimmune, polyglandular nature&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;APECED&lt;/td&gt;
      &lt;td&gt;Full clinical spectrum including ectodermal defects&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;The acronym may be awkward, but it reflects an important truth: this disorder is more than just endocrine autoimmunity.&lt;/p&gt;

&lt;h3 id=&quot;a-global-clue-founder-populations-1990s&quot;&gt;A Global Clue: Founder Populations (1990s)&lt;/h3&gt;

&lt;p&gt;One of the strongest hints that APS-1 was caused by a single gene came from geography. Worldwide, the disease was extremely rare, but it appeared much more often in a few genetically isolated populations: Finland, Sardinia, and Iranian Jewish communities. This pattern was classic for a founder effect, a mutation arising in a small ancestral group and becoming common over generations. Before whole-genome sequencing existed, geography itself functioned as a genetic tool. Population history narrowed the search space long before molecular methods could.&lt;/p&gt;

&lt;p&gt;Israeli geneticist John (Yehuda) Zlotogora reported that many Iranian Jewish cases shared the same founder mutation, suggesting a single ancestral origin. In parallel, clinicians like Piero Betterle in Italy and Eystein Husebye in Norway broadened the clinical map across Europe, showing that APS-1 looked remarkably consistent in its core features while varying in which organs were affected and when.&lt;/p&gt;

&lt;p&gt;By the early 1990s, the consensus was clear. APS-1 was not random, and it was not multifactorial in the usual sense. It looked like a single recessive gene disorder. The hunt for that gene was on.&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;gene-hunters-and-the-aire-of-discovery-mid-1990s&quot;&gt;Gene Hunters and the AIRE of Discovery (mid-1990s)&lt;/h2&gt;

&lt;p&gt;That hunt accelerated in the mid-1990s, when Finnish geneticist Leena Peltonen-Palotie and collaborators used large family pedigrees to map APS-1 to chromosome 21q22.3. Peltonen-Palotie was known for treating Finland’s population history not as a limitation, but as an experimental advantage. At a time when many geneticists favored large, diverse cohorts, she argued that bottlenecks and founder effects could turn rare diseases into solvable problems.&lt;/p&gt;

&lt;p&gt;APS-1 became one of her signature successes. In 1997, two independent teams arrived at the same answer almost simultaneously. A Finnish-German consortium associated with Peltonen’s work and an international NIH-backed group led by Kentaro Nagamine identified the same previously unknown gene whose disruption explained the syndrome.&lt;/p&gt;

&lt;p&gt;In the competitive climate of 1990s human genetics, simultaneous discovery often led to disputes over priority. In this case, it did the opposite. Two independent paths converged on the same gene, removing doubt and quickly unifying the field around a single name: AIRE, short for Autoimmune Regulator.&lt;/p&gt;

&lt;p&gt;From the start, AIRE’s protein sequence looked like it belonged in the nucleus rather than on the cell surface. It contained two PHD zinc finger domains, proline-rich regions, and LXXLL motifs associated with transcriptional regulation. PHD stands for Plant Homeodomain, a zinc-binding motif found in many chromatin-associated regulators. This architecture suggested that AIRE controlled other genes rather than acting as a structural or signaling protein. It looked like a master switch, exactly the kind of mechanism one might expect in a disorder where immune tolerance collapses across multiple organs.&lt;/p&gt;

&lt;p&gt;APS-1 also made history. It became one of the first clear examples of a single-gene autoimmune disease discovered outside the HLA region.&lt;/p&gt;

&lt;h3 id=&quot;founder-mutations-identified-late-1990s&quot;&gt;Founder Mutations Identified (late 1990s)&lt;/h3&gt;

&lt;p&gt;The gene discovery immediately explained the founder populations. Different groups tended to carry different recurring mutations:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Population&lt;/th&gt;
      &lt;th&gt;Common Mutation&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Finns&lt;/td&gt;
      &lt;td&gt;R257X (70-90% homozygous)&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Sardinians&lt;/td&gt;
      &lt;td&gt;R139X (~82% of alleles)&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Iranian Jews&lt;/td&gt;
      &lt;td&gt;Y85C (missense, unstable protein)&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Each population carried its own genetic signature, but the conclusion was the same. When AIRE failed, self-tolerance failed with it.&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;aires-job-teaching-the-immune-system-self-2001-to-2002&quot;&gt;AIRE’s Job: Teaching the Immune System ‘Self’ (2001 to 2002)&lt;/h2&gt;

&lt;p&gt;Finding AIRE solved the genetic mystery, but it raised a deeper question: what did AIRE actually do? The answer came from work focused on an organ most people rarely think about: the thymus, where developing T cells learn the difference between self and not-self. In 2001, immunologist Bruno Kyewski described something that initially sounded almost implausible. Certain thymic cells expressed an eclectic mix of genes normally restricted to entirely different organs.&lt;/p&gt;

&lt;p&gt;At the time, tissue specificity was considered almost sacred. The idea that a thymic cell might express insulin, eye proteins, or liver enzymes sounded like experimental noise. Kyewski argued otherwise: he proposed that the thymus deliberately created an internal self portrait by turning on tissue-specific genes in the wrong place, for the right reason. These specialized cells, medullary thymic epithelial cells, used this so-called promiscuous gene expression to present a broad sampling of self-antigens to developing T cells during training.&lt;/p&gt;

&lt;p&gt;The decisive tests came in 2002. Two groups independently created Aire-deficient mice. One of the most influential efforts came from the lab of Diane Mathis and Christophe Benoist, a scientific partnership known for turning abstract immunological ideas into clean experiments. Their team, with Mark Anderson playing a central role, removed Aire and watched a familiar pattern unfold: multi-organ autoimmunity, high autoantibody levels, and immune attacks on tissues that should have been protected.&lt;/p&gt;

&lt;p&gt;What mattered most was not just that the mice became sick, but how. Without Aire, many tissue-specific antigens were no longer expressed in the thymus. The immune system graduated with gaps in its education. Self-reactive T cells that should have been eliminated survived, escaped into circulation, and later attacked organs that had never appeared in the thymic lesson plan. Parallel work linked to Peltonen’s collaborations confirmed that Aire’s critical action occurred within the thymic environment itself.&lt;/p&gt;

&lt;p&gt;Mark Anderson’s later work added another twist. APS-1 patients often produce autoantibodies against type I interferons, a finding that helped explain why infection susceptibility and autoimmunity can coexist in the same syndrome. It also provided clinicians with a distinctive diagnostic signal, connecting bench discoveries back to patient care.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/img/20260101-overview.jpg&quot; alt=&quot;Overview of the AIRE story&quot; /&gt;&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;epilogue-from-rare-disease-to-foundational-immunology&quot;&gt;Epilogue: From Rare Disease to Foundational Immunology&lt;/h2&gt;

&lt;p&gt;We now know that as part of the central tolerance mechanism, developing T cells are constantly tested in the thymus. If a T cell reacts too strongly to self, it is removed, a process known as negative selection. AIRE’s key role is to help thymic cells display a wide range of self-antigens, including proteins normally found only in peripheral tissues. When AIRE is missing or broken, the immune system does not get the full picture of self, and dangerous self-reactive cells can slip through.&lt;/p&gt;

&lt;p&gt;That understanding feels straightforward today, but much of it once sounded strange. The idea that the thymus would deliberately violate tissue specificity, or that a single gene could orchestrate such a broad tolerance program, ran against intuition. Autoimmunity was often framed as irreducibly complex. AIRE made a simpler and more unsettling point. Sometimes one broken part is enough to destabilize the whole system.&lt;/p&gt;

&lt;p&gt;Even the gene hunt itself reflects a moment in time. In the 1990s, identifying a disease gene was closer to detective work than pipeline biology. Founder populations, family pedigrees, and careful mapping through anonymous stretches of DNA did the heavy lifting. When two independent teams converged on AIRE in 1997, it felt less like routine progress and more like resolution.&lt;/p&gt;

&lt;p&gt;Today, APS-1/APECED remains rare, and treatment is still largely supportive, hormone replacement for endocrine failure and antifungals for chronic infections. But the discovery of AIRE transformed the condition from a puzzling clinical triad into a foundational lesson about the immune system. It showed that immune tolerance depends on a real biological process, in a real place, driven by specific genes. It also changed how the thymus is viewed, not as a passive organ you outgrow, but as an active classroom with a surprisingly broad syllabus.&lt;/p&gt;

&lt;p&gt;In that sense, APS-1 did something remarkable. It took a rare disorder and used it to explain something universal: how the immune system learns when to fight, and when to leave home alone. And it reminds us that what feels obvious in hindsight is often the result of a few brave ideas that once sounded strange.&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;references&quot;&gt;References&lt;/h2&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;a href=&quot;https://academic.oup.com/jes/article/5/12/bvab151/6374539&quot;&gt;Hypoadrenalism as the Single Presentation of Autoimmune Polyglandular Syndrome Type 1 (Journal of the Endocrine Society, Oxford Academic)&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://clinicalimagingscience.org/autoimmune-polyglandular-syndrome-type-1/&quot;&gt;Autoimmune Polyglandular Syndrome Type 1 (Journal of Clinical Imaging Science)&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://en.wikipedia.org/wiki/Autoimmune_polyendocrine_syndrome_type_1&quot;&gt;Autoimmune polyendocrine syndrome type 1 (Wikipedia)&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://primaryimmune.org/understanding-primary-immunodeficiency/types-of-pi/apecedaps-1&quot;&gt;APECED/APS-1 (Immune Deficiency Foundation)&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://link.springer.com/article/10.1186/s12881-019-0870-3&quot;&gt;Autoimmune Polyglandular Syndrome Type 1: a case report (BMC Medical Genetics)&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.nature.com/articles/ng1297-393?error=cookies_not_supported&amp;amp;code=d8ddd219-ebf6-4703-8228-7bfb266f5e0e&quot;&gt;Positional cloning of the APECED gene (Nature Genetics)&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.nature.com/articles/ng1297-399?error=cookies_not_supported&amp;amp;code=ea42cb1d-997a-405c-9039-ae0c8db1ecf1&quot;&gt;An autoimmune disease, APECED, caused by mutations in a novel gene featuring two PHD-type zinc-finger domains (Nature Genetics)&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.ncbi.nlm.nih.gov/clinvar/RCV000169129/&quot;&gt;NM_000383.4(AIRE):c.415C&amp;gt;T (p.Arg139Ter) AND Polyglandular autoimmune syndrome, type 1 (ClinVar, NCBI)&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://pubmed.ncbi.nlm.nih.gov/12376594/&quot;&gt;Projection of an immunological self shadow within the thymus by the aire protein (PubMed)&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://pubmed.ncbi.nlm.nih.gov/11600886/&quot;&gt;Promiscuous gene expression in medullary thymic epithelial cells mirrors the peripheral self (PubMed)&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://academic.oup.com/hmg/article-abstract/11/4/397/550346&quot;&gt;Aire deficient mice develop multiple features of APECED phenotype (Human Molecular Genetics, Oxford Academic)&lt;/a&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;hr /&gt;

&lt;table&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;a href=&quot;/ai/vibe/research/rag1/rag2/vdj/tcr/bcr/2025/12/29/gene-stories-tcr-rag-rag1-rag2.html&quot;&gt;Part I&lt;/a&gt;&lt;/td&gt;
      &lt;td&gt;Part II (you are reading this)&lt;/td&gt;
      &lt;td&gt;&lt;a href=&quot;/ai/vibe/research/fas/fasl/cd95/apo1/tcr/bcr/2026/01/05/gene-stories-tcr-fas-cd95.html&quot;&gt;Part III&lt;/a&gt;&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;
</content>
 </entry>
 
 <entry>
   <title>RAG1 and RAG2: Discovery, Mechanism, and Evolution of the V(D)J Recombinase</title>
   <link href="https://ergoso.me/ai/vibe/research/rag1/rag2/vdj/tcr/bcr/2025/12/29/gene-stories-tcr-rag-rag1-rag2.html"/>
   <updated>2025-12-29T00:00:01+00:00</updated>
   <id>https://ergoso.me/ai/vibe/research/rag1/rag2/vdj/tcr/bcr/2025/12/29/gene-stories-tcr-rag-rag1-rag2</id>
   <content type="html">&lt;table&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Part I (you are reading this)&lt;/td&gt;
      &lt;td&gt;&lt;a href=&quot;/ai/vibe/research/aire/rag2/vdj/tcr/bcr/2026/01/01/gene-stories-tcr-aire.html&quot;&gt;Part II&lt;/a&gt;&lt;/td&gt;
      &lt;td&gt;&lt;a href=&quot;/ai/vibe/research/fas/fasl/cd95/apo1/tcr/bcr/2026/01/05/gene-stories-tcr-fas-cd95.html&quot;&gt;Part III&lt;/a&gt;&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;The adaptive immune system depends on a controlled act of genomic violence. To recognize an almost infinite variety of pathogens, developing lymphocytes deliberately break and rejoin their own DNA, assembling antigen receptor genes from modular fragments. For decades, immunologists understood the outcome of this process but not its cause: what enzyme could cut the genome so precisely, and why did it act only in immune cells?&lt;/p&gt;

&lt;p&gt;The answer emerged at the end of the 1980s with the discovery of RAG1 and RAG2, two genes that together form the molecular “scissors” of V(D)J recombination. Their identification not only solved a long-standing mystery in immunology but also revealed that the origins of adaptive immunity lie in an ancient, repurposed genetic element. This account follows the scientific hunt for the RAG genes and the far-reaching consequences of their discovery.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/img/20251229-timeline.jpg&quot; alt=&quot;RAG story timeline&quot; /&gt;&lt;/p&gt;

&lt;h2 id=&quot;the-hunt-for-the-immunological-scissors-1970s1980s&quot;&gt;The Hunt for the Immunological “Scissors” (1970s–1980s)&lt;/h2&gt;

&lt;p&gt;In the 1970s, Susumu Tonegawa (a Japanese molecular biologist trained in Tokyo and California) stunned the scientific world by showing that antibody genes rearrange in B cells, which was a finding that earned him the 1987 Nobel Prize. This process, now called V(D)J recombination, explained how a finite genome could produce the vast diversity of antibodies and T-cell receptors. Yet a mystery remained: what molecular “scissors” cut and rejoin DNA to make these receptor genes? Through the 1980s, immunologists imagined a dedicated V(D)J recombinase that was selectively expressed in lymphocytes was behind this phenomenon. The race was on to find it. Remarkably, Tonegawa did not participate in this race for too long as after securing his place in immunology history, he switched fields to neuroscience, applying his scientific boldness to memory and learning research.&lt;/p&gt;

&lt;p&gt;By the late 1980s, one approach to find the elusive recombinase was brute-force genetics. Among those intrigued was David Baltimore, already famous for co-discovering reverse transcriptase (Nobel Prize, 1975) and now pivoting to immunology. Baltimore, at MIT’s Whitehead Institute, suspected that transferring the right gene into non-immune cells might confer the ability to perform V(D)J recombination. It was a high-risk idea—&lt;em&gt;like finding a needle in a genomic haystack&lt;/em&gt;—but worth a shot.&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;cloning-rag1-and-rag2-breathing-life-into-fibroblasts-19881990&quot;&gt;Cloning RAG1 and RAG2: Breathing Life into Fibroblasts (1988–1990)&lt;/h2&gt;

&lt;p&gt;In 1988, Baltimore’s team (most notably graduate students David Schatz and Marjorie Oettinger) developed an ingenious assay. They engineered a reporter plasmid containing artificial gene segments and RSSs (Recombination Signal Sequences) that would produce a detectable signal &lt;em&gt;only if&lt;/em&gt; V(D)J recombination occurred. This reporter was transfected into NIH-3T3 fibroblasts, cells normally incapable of gene rearrangement, along with immune-cell DNA.&lt;/p&gt;

&lt;p&gt;Through serial genomic transfections, they isolated a genomic locus that conferred recombination activity. In late 1989, they identified a gene encoding a 1040–amino acid protein that activated V(D)J recombination in fibroblasts. They named it RAG1 (Recombination Activating Gene 1). RAG1 expression was restricted to lymphoid tissues—exactly where the scissors should be active.&lt;/p&gt;

&lt;p&gt;This was a great accomplishment but a piece of the puzzle was still missing: RAG1 alone was not sufficient for maximal activity, potentially hinting at a missing partner that was yet to be discovered. 2 years after their discovery of RAG1, in June 1990, Oettinger and Schatz reported RAG2, a gene adjacent to RAG1 that synergizes with it to produce full recombination activity. Interestingly, RAG2 was lacking obvious enzymatic motifs but without RAG2, RAG1 was a dull knife. Only together, RAG1 and RAG2 reconstituted the V(D)J recombinase at expected levels.&lt;/p&gt;

&lt;p&gt;The molecular scissors were a two-part machine. After these breakthough findings, then-graduate-students David Schatz and Marjorie Oettinger eventually became leaders thmselves and continued working on V(DJ) recombination and the RAG proteins. David Baltimore later reflected that identifying RAG proteins was among his lab’s proudest achievements.&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;proof-by-knockout-no-rag-no-immune-system-1992&quot;&gt;Proof by Knockout: No RAG, No Immune System (1992)&lt;/h2&gt;

&lt;p&gt;Discovery in vitro was only the beginning. In 1992, two groups independently generated RAG knockout mice. Published back-to-back in &lt;em&gt;Cell&lt;/em&gt;, the results were unequivocal.&lt;/p&gt;

&lt;p&gt;Peter Mombaerts (of Tonewaga lab) showed that in RAG1-KO mice, there were no mature B or T lymphocytes; thymus and lymph nodes were severely underdeveloped; immunoglobulin and TCR genes remained unrearranged; and interestingly the mice phenotype resembled a known human disease (more on this later). In addition to these, with this knockout mouse, Mombaerts showed that despite low-level brain expression reports of RAG1, RAG1-KO mice were neurologically normal. Like his mentor, Mombaerts also switched to neurogenetics and eventually become a leader in olfactory system research.&lt;/p&gt;

&lt;p&gt;In parallel, Yuko Shinkai (of Frederick Alt lab) showed that in RAG2-KO mice, early lymphoid progenitors were present; V(D)J recombination initiation was completely failing; there was no detectable rearranged antigen receptor genes; and reintroduction of RAG2 rescued recombination in cultured cells. Shinkai, later, followed RAG2’s link to histone modifications and built his careers around epigenetics by studying chromatin regulation.&lt;/p&gt;

&lt;p&gt;Together, these two amazing knockouts and their characterization proved that RAG1 and RAG2 were absolutely required for adaptive immunity. The details of how this pair of proteins accomplished recombination followed later on.&lt;/p&gt;

&lt;h2 id=&quot;rag-mechanism-and-evolution-jumping-genes-to-immune-genes&quot;&gt;RAG Mechanism and Evolution: Jumping Genes to Immune Genes&lt;/h2&gt;

&lt;p&gt;Biochemical studies in the mid-1990s revealed that: RAG1 contained the catalytic DDE motif, typical of transposases; RAG2 acted as an essential cofactor; and core RAG1/2 proteins alone could catalyze recombination in vitro. Then came the stunning discovery of RAG proteins’ ability to mediate DNA transposition under artificial conditions, which led to the transposon hypothesis that RAG genes originated from an ancient mobile genetic element. It wasn’t until 2016 that the discovery of ProtoRAG in amphioxus provided compelling evolutionary evidence to this claim.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/img/20251229-overview.jpg&quot; alt=&quot;Overview of the RAG story&quot; /&gt;&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;epilogue-legacy-of-the-rag-story&quot;&gt;Epilogue: Legacy of the RAG Story&lt;/h2&gt;

&lt;p&gt;The discovery of RAG1 and RAG2 resolved a decades-old mystery and reshaped immunology, genetics, evolution, and medicine. What began as a search for DNA scissors revealed the molecular basis of immune diversity, the evolutionary domestication of transposons, and the genetic roots of immunodeficiency and autoimmunity. We now know that severe Combined Immunodeficiency (SCID) is a human disease that is caused by null RAG mutations. When these RAG mutations are hypomorphic, they then lead to Omenn syndrome where patients show very limited T cell development and autoimmune-like inflammation. In some cases, bone marrow transplantation from healthy donors can cure these conditions by restoring functional RAG genes.&lt;/p&gt;

&lt;p&gt;From clever reporter assays to definitive knockout models, the RAG story exemplifies how bold ideas, technical rigor, and human curiosity converge to answer fundamental biological questions. It is very interesting that two unassuming genes changed how we understand the immune system and life’s capacity for innovation.&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;references&quot;&gt;References&lt;/h2&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.cell.com/cell/fulltext/S0092-8674%2816%2930789-9&quot;&gt;Evidence of G.O.D.’s Miracle: Unearthing a RAG Transposon&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://link.springer.com/article/10.1385/IR:23:1:23&quot;&gt;RAG1 and RAG2 in V(D)J Recombination and Transposition&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.frontiersin.org/articles/10.3389/fimmu.2013.00110/full&quot;&gt;A Novel Quantitative Fluorescent Reporter Assay for RAG Targets and RAG Activity&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://www.nobelprize.org/prizes/medicine/1975/baltimore/biographical/&quot;&gt;David Baltimore – Biographical&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://authors.library.caltech.edu/records/vpxr8-1xr41/latest&quot;&gt;The V(D)J Recombination Activating Gene, RAG-1&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://link.springer.com/article/10.1093/emboj/16.10.2656&quot;&gt;Role of Recombination Signal Sequences in Formation of Coding and Signal Joints&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://pubmed.ncbi.nlm.nih.gov/1547488/&quot;&gt;RAG-1–Deficient Mice Have No Mature B and T Lymphocytes&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://pubmed.ncbi.nlm.nih.gov/1547487/&quot;&gt;RAG-2–Deficient Mice Lack Mature Lymphocytes Owing to Inability to Initiate V(D)J Rearrangement&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://ritaallen.org/all-scholars/peter-mombaerts/&quot;&gt;Peter Mombaerts – Rita Allen Foundation Profile&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://shinkai.riken.jp/en/member/shinkai.html&quot;&gt;Cellular Memory Laboratory – Yoichi Shinkai&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;https://pmc.ncbi.nlm.nih.gov/articles/PMC31291/&quot;&gt;The RAG Proteins in V(D)J Recombination: More Than Just a Nuclease&lt;/a&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;hr /&gt;

&lt;table&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Part I (you are reading this)&lt;/td&gt;
      &lt;td&gt;&lt;a href=&quot;/ai/vibe/research/aire/rag2/vdj/tcr/bcr/2026/01/01/gene-stories-tcr-aire.html&quot;&gt;Part II&lt;/a&gt;&lt;/td&gt;
      &lt;td&gt;&lt;a href=&quot;/ai/vibe/research/fas/fasl/cd95/apo1/tcr/bcr/2026/01/05/gene-stories-tcr-fas-cd95.html&quot;&gt;Part III&lt;/a&gt;&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;
</content>
 </entry>
 
 <entry>
   <title>p53: The gene that took a decade to become itself</title>
   <link href="https://ergoso.me/ai/vibe/research/tp53/p53/2025/12/25/gene-stories-p53-tp53.html"/>
   <updated>2025-12-25T00:00:01+00:00</updated>
   <id>https://ergoso.me/ai/vibe/research/tp53/p53/2025/12/25/gene-stories-p53-tp53</id>
   <content type="html">&lt;p&gt;If you have ever heard &lt;strong&gt;p53&lt;/strong&gt; (a protein encoded by the &lt;strong&gt;TP53&lt;/strong&gt; gene) called &lt;em&gt;“the guardian of the genome,”&lt;/em&gt; you might picture a wise molecular sentry that patrols DNA, stops cell cycles, and flips apoptosis switches like a seasoned air-traffic controller. But p53 didn’t make its debut as a heroic gene. p53, actually, began as a mysterious band on a gel as a stubborn ~53 kDa shadow that kept showing up whenever scientists poked at cancer-causing viruses. And for years, the field argued about what it meant; because the tools of the time could show you &lt;em&gt;something was there&lt;/em&gt;, but not &lt;em&gt;what it really was&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;This is the story of how p53 went from being a molecular hitchhiker to an accidental “oncogene” then to the most famous tumor suppressor on Earth; and how a few very human scientific decisions (and missteps) shaped everything that followed.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/img/20251225-timeline.jpg&quot; alt=&quot;p53 story timeline&quot; /&gt;&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;a-fishing-expedition-in-1979&quot;&gt;&lt;strong&gt;A fishing expedition in 1979&lt;/strong&gt;&lt;/h2&gt;

&lt;p&gt;In London, 1979, a cancer virus called SV40 was the hot object. Researchers were trying to understand how viral proteins hijack cells. One of those researchers was Sir David Lane, the co-discoverer who later &lt;em&gt;coined the phrase&lt;/em&gt; “guardian of the genome” and was knighted in 2000 for contributions to cancer research. He also co-edited the bench-classic &lt;em&gt;Antibodies: A Laboratory Manual&lt;/em&gt; with Ed Harlow—one of those books that ends up permanently warped from humidity on lab shelves.&lt;/p&gt;

&lt;p&gt;Sir Lane and his colleauge, Lionel Crawford, weren’t looking for p53 per se; they were doing what the field often does best: following the weird thing that refuses to go away. Using immunoprecipitation with anti–T-antigen antibodies (Lane later called it essentially a “fishing expedition”), they pulled down SV40 large T antigen, and with it, a mysterious ~53 kDa host protein that seemed to bind T antigen over and over again.&lt;/p&gt;

&lt;p&gt;Even in that first moment, there’s something delightful: Lane and Crawford reasoned the protein had to be host-derived, because SV40’s tiny genome couldn’t encode for the gene behind that extra big band. And in their first Nature paper, they floated a bold idea that this host protein might normally regulate growth control. That band would become one of the most consequential in modern biology.&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;meanwhile-in-princeton-another-gel-another-band-same-ghost&quot;&gt;&lt;strong&gt;Meanwhile in Princeton: another gel, another band, same ghost&lt;/strong&gt;&lt;/h2&gt;

&lt;p&gt;A few months later, another group sees the same kind of thing. This group of researchers was led by  Arnold J. Levine who was born in Brooklyn (1939). Levine, later, became the president of Rockefeller University. Rockefeller’s own announcement notes he became its 8th president in 1998, after years building molecular biology at Princeton. He would act as one of the central characters in the p53 saga.&lt;/p&gt;

&lt;p&gt;Daniel Linzer, a graduate student in Levine’s group, isolated a 54 kDa protein in SV40-transformed cells using sera from tumor-bearing animals. Peptide mapping suggested it was distinct from viral antigens, reinforcing that this too was a host protein induced during transformation. After graduating with a seminal thesis that included the discovery of p53, Dan Linzer later became a major academic leader and research administrator.&lt;/p&gt;

&lt;p&gt;At roughly the same time, multiple other teams detected similar ~53–55 kDa proteins in transformed cells. By late 1979, it was clear: different groups, different assays, same molecular “something.” . That “something” was later named as &lt;strong&gt;p53&lt;/strong&gt; (for its apparent molecular weight), thanks to Lloyd Old’s immunology group. Beside being a legendary tumor immunologist, Lloyd Old had a deep love of music and he was often described as an accomplished violinist. In that sense, he is not unique: many of the scientists who built modern cancer biology had entire parallel lives that involved music, art, and languages.&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;toolbox-check-what-1979-could-and-couldnt-do&quot;&gt;&lt;strong&gt;Toolbox check: what 1979 could (and couldn’t) do&lt;/strong&gt;&lt;/h2&gt;

&lt;p&gt;To appreciate why p53’s story gets messy, it helps to remember a few things. In 1979, p53 was basically &lt;em&gt;a recurring band on SDS-PAGE&lt;/em&gt;. Back then, scientists could detect peptides by immunoprecipitation and gels; they could compare peptide maps to confirm it’s the same protein that is seen by others; you could infer it’s host-encoded or not; but, you couldn’t easily clone/sequence it on demand, or do fast mutational surveys across tumors.&lt;/p&gt;

&lt;p&gt;So, early p53 was characterized biochemically and immunologically, long before anyone knew its DNA sequence or true function. This matters, because when the gene finally &lt;em&gt;was&lt;/em&gt; cloned, the versions that are easiest to clone are often not the “normal” ones.&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;the-great-misunderstanding-earlymid-1980s&quot;&gt;&lt;strong&gt;The great misunderstanding (early–mid 1980s)&lt;/strong&gt;&lt;/h2&gt;

&lt;p&gt;By the early 1980s, scientists started asking: &lt;em&gt;When does p53 show up in normal cells?&lt;/em&gt; Reich and Levine showed p53 levels rise when quiescent cells were stimulated to re-enter the cell cycle, which was supporting an early interpretation of p53 as a growth-linked regulator. Even more dramatic was that Mercer and colleagues microinjected anti-p53 antibodies and found cells could fail to enter S-phase, which was another nudge toward the “p53 helps proliferation” story.&lt;/p&gt;

&lt;p&gt;And then came the cloning race.&lt;/p&gt;

&lt;h3 id=&quot;the-cloning-grind-and-a-near-abandonment-moment&quot;&gt;The cloning grind (and a near-abandonment moment)&lt;/h3&gt;

&lt;p&gt;Cloning TP53 in the early ’80s was laborious, failure-prone, and brutally slow. Moshe Oren, who was a key figure who initially helped push p53 into the “oncogene” narrative, then later helped overturn it with data. His arc is basically the plot twist in human form. Oren recalled repeated failures so discouraging that Levine briefly considered abandoning p53 around 1981.&lt;/p&gt;

&lt;p&gt;Eventually, cloning succeeded but there was still a catch: Tumor cells often overexpress p53. Those are the samples that jump out at you. Those are the samples you build libraries from. Those are the samples that give you clones. And (as the field would learn the hard way) many of those clones encoded mutant p53. So, p53 was labeled a proto-oncogene not because scientists were careless; but because they were doing the best possible science with the most available material and tools.&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;1989-the-year-the-story-flips&quot;&gt;&lt;strong&gt;1989: the year the story flips&lt;/strong&gt;&lt;/h2&gt;

&lt;p&gt;By the late 1980s, the cracks in the “p53 is an oncogene” story were getting impossible to ignore.&lt;/p&gt;

&lt;p&gt;Finlay, Hinds, and Levine noticed that a p53 cDNA from a normalish context (e.g., F9 embryonal carcinoma) didn’t behave like the tumor-derived clones in transformation assays. Phil Hinds discovered that many cloned p53 sequences differed by single amino acid substitutions. These were first thought to be polymorphisms but then soon recognized as mutations. That realization is one of my favorite scientific moments: not flashy, not cinematic. It was just the slow dawning that the “same gene” everybody has been studying wasn’t actually the same gene in everybody’s experiments.&lt;/p&gt;

&lt;p&gt;Then two kinds of evidence arrive like a one-two punch:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;strong&gt;Functional proof that wild-type p53 suppresses transformation&lt;/strong&gt;: Finlay et al. (Cell) showed wild-type p53 could suppress oncogene-driven transformation. Eliyahu et al. (PNAS) then confirmed similar suppression and even early hints of growth arrest.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Genetic proof that p53 is mutated in human tumors&lt;/strong&gt;: Baker, Fearon, Vogelstein, and others sequenced colorectal tumors and found frequent TP53 mutations often paired with 17p loss, which was the classic tumor suppressor “two-hit” logic.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;It is worth mentioning that Bert Vogelstein, another key figure in the p53 story, was a math major who also won an undergraduate award for Semitic languages and literature at Penn (yes, really). His genetic “sleuthing” approach helped shift p53 from cell-culture argument to human-cancer fact. So, he was yet another scientist who had diverse interests spanning more than the biology field.&lt;/p&gt;

&lt;p&gt;Within about a year after all these happened, p53 went from “maybe an oncogene” to “the most commonly mutated gene in cancer,” and the field’s collective interpretation snapped into a new shape.&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;when-p53-becomes-personal-li-fraumeni-1990&quot;&gt;&lt;strong&gt;When p53 becomes personal (Li-Fraumeni, 1990)&lt;/strong&gt;&lt;/h2&gt;

&lt;p&gt;If the tumor data convinced the scientists, Li-Fraumeni syndrome convinced everyone else. Two groups found germline TP53 mutations in Li-Fraumeni families, which suddenly connected the molecular story to a human one: families haunted by sarcomas, breast cancers, brain tumors, and early-onset malignancies…&lt;/p&gt;

&lt;p&gt;At this point, debates about p53 were not about assays but the whole narrative became a disease story.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/img/20251225-overview.jpg&quot; alt=&quot;Overview of the p53 story&quot; /&gt;&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;1992---guardian-of-the-genome-and-the-knockout-mice-that-sealed-it&quot;&gt;&lt;strong&gt;1992 - “Guardian of the genome” (and the knockout mice that sealed it)&lt;/strong&gt;&lt;/h2&gt;

&lt;p&gt;1992 delivered two milestone breakthroughs:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;strong&gt;Knockout mice: the brutal clarity of genetics&lt;/strong&gt;:  Donehower, Jacks, and others created p53 knockout mice. They developed normally, but were highly prone to spontaneous tumors early in life (often by 4–6 months). That finding is one of those rare science results that feels like a gavel strike: &lt;em&gt;case closed.&lt;/em&gt;&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;The phrase that stuck: “guardian of the genome”&lt;/strong&gt;: Lane wrote a short Nature commentary titled “p53, guardian of the genome,” crystallizing the idea of p53 as a sentinel that halts the cell cycle or triggers cell death in response to DNA damage. The 1992 paper is a tiny one-page piece with an outsized cultural footprint.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;With genetics and language finally aligned, p53’s role as a tumor suppressor was no longer a matter of debate but of detail. What followed was not a search for whether p53 mattered, but a deeper reckoning with how it worked and why cancer seemed so determined to disable it.&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;1993---superstardom-and-the-uneasy-next-chapter&quot;&gt;&lt;strong&gt;1993 - superstardom (and the uneasy next chapter)&lt;/strong&gt;&lt;/h2&gt;

&lt;p&gt;By 1993, p53’s reputation exploded. Science dubbed it the “Molecule of the Year.”  And then, as if the story needed one more twist, the field began to absorb a deeper complexity: mutant p53 wasn’t merely “loss of function”. Some mutants seemed to acquire gain-of-function behavior, a theme Levine and others pushed into the open by the early 1990s.&lt;/p&gt;

&lt;p&gt;The gene didn’t just have a heroic form and a broken form. It had multiple identities that are all context- and mutation-dependent. p53 was like a character whose motives changed depending on which chapter the reader was going through.&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;epilogue-p53-understood-at-last&quot;&gt;&lt;strong&gt;Epilogue: p53, understood at last&lt;/strong&gt;&lt;/h2&gt;

&lt;p&gt;What we now call &lt;strong&gt;TP53&lt;/strong&gt; encodes a transcription factor that acts less like a single-purpose brake and more like a cellular decision-maker. In healthy cells, p53 is kept low and fleeting; it is powerful enough that the cell actively restrains it. When genuine danger appears (e.g. DNA damage, oncogene activation, replication stress), p53 stabilizes and turns on gene programs that pause the cell cycle, enforce permanent arrest, or trigger cell death. Which outcome occurs depends on context: cell type, stress intensity, and the surrounding regulatory circuitry. The early confusion around p53 was not just technical. It was conceptual. Scientists were looking for one function, while p53’s biology is inherently conditional.&lt;/p&gt;

&lt;p&gt;Cancer’s relationship with p53 explains both its fame and its early mislabeling. TP53 is the most frequently mutated gene in human cancer, and most of those mutations are missense changes that produce a stable, altered protein rather than a simple loss. These mutants can disable normal p53 function, interfere with any remaining wild-type protein, and sometimes acquire new, tumor-promoting behaviors. In retrospect, it’s easy to see why p53 looked oncogenic in early assays: the field was unknowingly studying mutant versions while assuming they represented the gene’s true nature.&lt;/p&gt;

&lt;p&gt;Today, p53 stands as both a solved mystery and an ongoing challenge. We understand its structure, regulation, and role in tumor suppression with extraordinary depth, yet translating that knowledge into universal therapies remains difficult precisely because p53 is so context-dependent. The gene that began as a stubborn band on a gel ultimately taught the field something broader: biology rarely reveals its meaning all at once. Sometimes it takes a decade (sprinkled with wrong turns) for a gene to be fully understood.&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;references&quot;&gt;References&lt;/h2&gt;

&lt;ol&gt;
  &lt;li&gt;International Agency for Research on Cancer (IARC). &lt;em&gt;Professor Sir David Lane&lt;/em&gt;. Available at: &lt;a href=&quot;https://www.iarc.who.int/friends-of-iarc/professor-sir-david-lane/&quot;&gt;https://www.iarc.who.int/friends-of-iarc/professor-sir-david-lane/&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;Harlow E, Lane D (eds.). &lt;em&gt;Antibodies: A Laboratory Manual&lt;/em&gt;. Cold Spring Harbor Laboratory Press, New York; 1988. Available at: &lt;a href=&quot;https://www.cambridge.org/core/journals/genetics-research/article/antibodies-a-laboratory-manual-edited-by-ed-harlow-and-david-lane-cold-spring-harbor-cold-spring-harbor-laboratory-new-york-1988-726-pages-paper-5000-isbn-0-87969-314-2/4DCD2ECC6484BBE46541208D52F324CE&quot;&gt;https://www.cambridge.org/core/journals/genetics-research/article/antibodies-a-laboratory-manual-edited-by-ed-harlow-and-david-lane-cold-spring-harbor-cold-spring-harbor-laboratory-new-york-1988-726-pages-paper-5000-isbn-0-87969-314-2/4DCD2ECC6484BBE46541208D52F324CE&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;The Rockefeller University. &lt;em&gt;Arnold J. Levine Named President of Rockefeller University&lt;/em&gt;. Available at: &lt;a href=&quot;https://www.rockefeller.edu/news/4415-arnold-j-levine-named-president-of-rockefeller-university/&quot;&gt;https://www.rockefeller.edu/news/4415-arnold-j-levine-named-president-of-rockefeller-university/&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;The Rockefeller University. &lt;em&gt;Arnold J. Levine Becomes Eighth President of The Rockefeller University&lt;/em&gt;. Available at: &lt;a href=&quot;https://www.rockefeller.edu/news/4399-arnold-j-levine-becomes-eighth-president-of-the-rockefeller-university&quot;&gt;https://www.rockefeller.edu/news/4399-arnold-j-levine-becomes-eighth-president-of-the-rockefeller-university&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;University of Arizona, College of Science. &lt;em&gt;Dr. Dan Linzer&lt;/em&gt;. Available at: &lt;a href=&quot;https://science.arizona.edu/person/dr-dan-linzer&quot;&gt;https://science.arizona.edu/person/dr-dan-linzer&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;Old LJ. &lt;em&gt;Lloyd J. Old — a scientific concertmaster&lt;/em&gt;. Proc Natl Acad Sci U S A. Available at: &lt;a href=&quot;https://pmc.ncbi.nlm.nih.gov/articles/PMC3337005/&quot;&gt;https://pmc.ncbi.nlm.nih.gov/articles/PMC3337005/&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;Johns Hopkins Technology Ventures. &lt;em&gt;Bert Vogelstein, MD&lt;/em&gt;. Available at: &lt;a href=&quot;https://ventures.jhu.edu/jhtv-events/celebration-of-innovation-in-medicine-2025/bert-vogelstein-md/&quot;&gt;https://ventures.jhu.edu/jhtv-events/celebration-of-innovation-in-medicine-2025/bert-vogelstein-md/&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;Lane DP. &lt;em&gt;p53, guardian of the genome&lt;/em&gt;. &lt;em&gt;Nature&lt;/em&gt;. 1992;358:15–16. Available at: &lt;a href=&quot;https://www.nature.com/articles/358015a0&quot;&gt;https://www.nature.com/articles/358015a0&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;Wikipedia contributors. &lt;em&gt;Lionel Crawford&lt;/em&gt;. Wikipedia. Available at: &lt;a href=&quot;https://en.wikipedia.org/wiki/Lionel_Crawford&quot;&gt;https://en.wikipedia.org/wiki/Lionel_Crawford&lt;/a&gt;&lt;/li&gt;
&lt;/ol&gt;

</content>
 </entry>
 
 <entry>
   <title>Vibe researching: Making sense of DepMap's extreme responders via GPT 5 Pro</title>
   <link href="https://ergoso.me/ai/vibe/research/depmap/crispr/2025/12/08/vibe-researching-part-one-depmap-extreme-responder.html"/>
   <updated>2025-12-08T09:00:00+00:00</updated>
   <id>https://ergoso.me/ai/vibe/research/depmap/crispr/2025/12/08/vibe-researching-part-one-depmap-extreme-responder</id>
   <content type="html">&lt;blockquote&gt;
  &lt;p&gt;Akin to vibe coding, vibe researching is a modern approach to learning where you use LLMs to dive into a topic by prompting, exploring, and iterating—essentially ‘vibing’ your way through the research while the AI handles the deep reading and connects the dots for you.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Let’s imagine you are casually exploring DepMap and run into this strange and challenging profile:&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/img/20251208-extremeresponders-grs.jpg&quot; alt=&quot;SK-MES-1 is an extreme-responder for GRS&quot; /&gt;&lt;/p&gt;

&lt;p&gt;What is unusual here is that almost all cell lines show a neutral response when &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;GSR&lt;/code&gt; is depleted via CRISPR KO, yet this single cancer cell line, called &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;SK-MES-1&lt;/code&gt;, shows extreme sensitivity. The hard part is that this is an &lt;em&gt;n=1&lt;/em&gt; case, which limits what you can do with standard machine-learning comparisons (e.g., genomic differences across responders vs. non-responders). Because of that, these types of outliers are in the data are often ignored.&lt;/p&gt;

&lt;p&gt;I love digging into these phenotypes, and, in the past, I spent hours, days, sometimes weeks dissecting multimodal data and combing through the literature on a cell line or gene. This is especially fun when you actually have the time (e.g. as part of graduate studies) or if this is part of a formal early-discovery effort (i.e. you get paid for it). One such deep dive during my PhD eventually became &lt;a href=&quot;https://www.biorxiv.org/content/10.1101/005686v2&quot;&gt;a preprint on an obscure oncogenic DICER1 mutation seen across a handful of patients&lt;/a&gt;. However, most of these explorations end up revealing a sample or data issue. Maybe 2 out of 10 point toward a pathway but stop short of a clear explanation. And if I get lucky, 1 out of 10 leads to a meaningful reason behind the extreme result, only to be dismissed as non-generalizable or in need of experimental follow-up. So these deep dives rarely end up as publications, and &lt;a href=&quot;https://ergoso.me/cancer/tcga/mutations/mgam/2014/01/28/mgam.html&quot;&gt;this is exactly why I wanted to start this blog originally&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Back to &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;SK-MES-1&lt;/code&gt; and its extreme dependency on &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;GSR&lt;/code&gt; (Glutathione Disulfide Reductase): I had never encountered this cell line or gene before, so the biology felt like a blank slate. I suspected there was a specific genetic context—maybe a mutation, CNV, or fusion—hidden somewhere on this page:&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/img/20251208-extremeresponders-skmes1.jpg&quot; alt=&quot;SK-MES-1 profile on DepMap&quot; /&gt;&lt;/p&gt;

&lt;p&gt;I have been using and experimenting with deep-research utilities since their early days, but until recently we didn’t have a reliable API to run them at scale. Submitting one request at a time through a web interface—with strict rate limits—just wasn’t workable. Now, with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;gpt-5-pro&lt;/code&gt; exposed through OpenAI’s API (and alternatives like Parallel.io), I decided to build a prompt template. After some trial and error, I settled on this:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;A cancer cell line, called {cell_line_name}, is extremely sensitive to gene {gene} depletion via CRISPR-KO. Top 10 major dependencies for this cell line are: {top10deps}.&lt;/p&gt;

  &lt;p&gt;Here are the mutations identified in this cell line:
{mutation_table}&lt;/p&gt;

  &lt;p&gt;Based on these, can you help me understand why {cell_line_name} might be super sensitive to {gene} depletion? Use markdown when you are answering and make sure you keep the references in.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;To my surprise, it took less than three minutes for GPT-5-Thinking to identify the synthetic-lethality relationship out of the 426 mutations included:&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/img/20251208-extremeresponders-example_response.jpg&quot; alt=&quot;Reason why SK-MES-1 is highly dependent on GRS&quot; /&gt;&lt;/p&gt;

&lt;p&gt;and here is the part of the response that goes into detail:&lt;/p&gt;

&lt;blockquote&gt;

  &lt;p&gt;…&lt;/p&gt;

  &lt;h3 id=&quot;the-key-hit-txnrd1-splice-acceptor-variant&quot;&gt;1. The key hit: TXNRD1 splice-acceptor variant&lt;/h3&gt;

  &lt;p&gt;From your list:&lt;/p&gt;

  &lt;blockquote&gt;
    &lt;p&gt;&lt;strong&gt;TXNRD1 – SNV, splice_acceptor_variant&lt;/strong&gt;&lt;/p&gt;
  &lt;/blockquote&gt;

  &lt;p&gt;TXNRD1 encodes &lt;strong&gt;thioredoxin reductase 1&lt;/strong&gt;, the central reductase for the cytosolic thioredoxin system. Together with GSR, it makes up one of two &lt;strong&gt;major, partially redundant disulfide-reducing antioxidant systems&lt;/strong&gt;:&lt;/p&gt;

  &lt;ul&gt;
    &lt;li&gt;&lt;strong&gt;GSR&lt;/strong&gt; → keeps &lt;strong&gt;glutathione&lt;/strong&gt; (GSSG → 2 GSH) reduced&lt;/li&gt;
    &lt;li&gt;&lt;strong&gt;TXNRD1&lt;/strong&gt; → keeps &lt;strong&gt;thioredoxin&lt;/strong&gt; reduced (TXN–S₂ → TXN–(SH)₂)&lt;/li&gt;
  &lt;/ul&gt;

  &lt;p&gt;These two pathways are Nrf2-regulated and function as &lt;strong&gt;parallel antioxidant systems&lt;/strong&gt; in mammalian cells (&lt;a href=&quot;https://www.sciencedirect.com/science/article/pii/S2213231725001661?utm_source=chatgpt.com&quot;&gt;Distinct roles for the thioredoxin and glutathione antioxidant systems …&lt;/a&gt;).
Multiple studies have emphasized that &lt;strong&gt;TXNRD1 and GSR are functionally paired disulfide reductases&lt;/strong&gt;, and that simultaneous inhibition of both is catastrophic for tumor cells (&lt;a href=&quot;https://www.sciencedirect.com/science/article/pii/S2211124722000596?utm_source=chatgpt.com&quot;&gt;Characterization of a small molecule inhibitor of disulfide reductases …&lt;/a&gt;).&lt;/p&gt;

  &lt;p&gt;Key points:&lt;/p&gt;

  &lt;ul&gt;
    &lt;li&gt;&lt;strong&gt;Splice-acceptor variants&lt;/strong&gt; at canonical intron–exon boundaries are usually &lt;strong&gt;loss-of-function&lt;/strong&gt; (exon skipping / intron retention → frameshift → NMD).&lt;/li&gt;
    &lt;li&gt;If this TXNRD1 variant affects the main coding isoform, you effectively &lt;strong&gt;cripple the thioredoxin system&lt;/strong&gt; in that cell line.&lt;/li&gt;
    &lt;li&gt;In that context, the cell is forced to lean heavily on &lt;strong&gt;GSR + GSH&lt;/strong&gt; for redox homeostasis.&lt;/li&gt;
  &lt;/ul&gt;

  &lt;p&gt;There is good in vivo and in vitro evidence that:&lt;/p&gt;

  &lt;ul&gt;
    &lt;li&gt;When &lt;strong&gt;glutathione is depleted&lt;/strong&gt;, cells up-regulate or depend more on TXNRD1, and &lt;strong&gt;vice versa&lt;/strong&gt; – if TXNRD1 is lost, glutathione pathways become more critical (&lt;a href=&quot;https://www.sciencedirect.com/science/article/pii/S2352345X25001687?utm_source=chatgpt.com&quot;&gt;A Redundant System of Thioredoxin and Glutathione Is Essential for …&lt;/a&gt;).&lt;/li&gt;
    &lt;li&gt;Combined inhibition or loss of both &lt;strong&gt;GSR and TXNRD1&lt;/strong&gt; drives strong oxidative stress and cancer cell death (synthetic-lethal-like behavior) (&lt;a href=&quot;https://aacrjournals.org/cancerres/article/79/1/125/633217/Inhibition-of-Thioredoxin-Thioredoxin-Reductase?utm_source=chatgpt.com&quot;&gt;Inhibition of Thioredoxin/Thioredoxin Reductase Induces Synthetic …&lt;/a&gt;).&lt;/li&gt;
  &lt;/ul&gt;

  &lt;p&gt;…&lt;/p&gt;

&lt;/blockquote&gt;

&lt;p&gt;As you can see, a thinking model with light search capabilities was able to figure this out with almost no hand-holding. Doing this manually would have taken me days to track down the weak link and build the biological case. So the obvious next question: &lt;strong&gt;can this scale?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;To test it, &lt;a href=&quot;https://github.com/armish/vibe-researching/blob/main/depmap-extreme_responders/Identify%20extreme%20responders.ipynb&quot;&gt;I first created a simple method to identify extreme responders in DepMap&lt;/a&gt;. Here are the top 20:&lt;/p&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Cell line&lt;/th&gt;
      &lt;th&gt;Model ID&lt;/th&gt;
      &lt;th&gt;Gene&lt;/th&gt;
      &lt;th&gt;Gene CNV&lt;/th&gt;
      &lt;th&gt;Chronos&lt;/th&gt;
      &lt;th&gt;Mean Chronos&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;SKMES1&lt;/td&gt;
      &lt;td&gt;ACH-000665&lt;/td&gt;
      &lt;td&gt;GSR&lt;/td&gt;
      &lt;td&gt;0.689901355&lt;/td&gt;
      &lt;td&gt;-2.805081015&lt;/td&gt;
      &lt;td&gt;0.073210444&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;NCIH209&lt;/td&gt;
      &lt;td&gt;ACH-000290&lt;/td&gt;
      &lt;td&gt;SOX1&lt;/td&gt;
      &lt;td&gt;0.900322013&lt;/td&gt;
      &lt;td&gt;-2.330362872&lt;/td&gt;
      &lt;td&gt;0.107062906&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;HCC1395&lt;/td&gt;
      &lt;td&gt;ACH-000699&lt;/td&gt;
      &lt;td&gt;TXNL1&lt;/td&gt;
      &lt;td&gt;2.231122922&lt;/td&gt;
      &lt;td&gt;-2.212720497&lt;/td&gt;
      &lt;td&gt;-0.1103986&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;TF1&lt;/td&gt;
      &lt;td&gt;ACH-000387&lt;/td&gt;
      &lt;td&gt;ETV6&lt;/td&gt;
      &lt;td&gt;0.91171322&lt;/td&gt;
      &lt;td&gt;-2.164286582&lt;/td&gt;
      &lt;td&gt;-0.021010812&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;OS384&lt;/td&gt;
      &lt;td&gt;ACH-002834&lt;/td&gt;
      &lt;td&gt;SNAI2&lt;/td&gt;
      &lt;td&gt;1.434246503&lt;/td&gt;
      &lt;td&gt;-2.113822951&lt;/td&gt;
      &lt;td&gt;-0.112583652&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;MM100511&lt;/td&gt;
      &lt;td&gt;ACH-002928&lt;/td&gt;
      &lt;td&gt;BRAP&lt;/td&gt;
      &lt;td&gt;0.747675415&lt;/td&gt;
      &lt;td&gt;-1.978243035&lt;/td&gt;
      &lt;td&gt;-0.255199889&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;EKVX&lt;/td&gt;
      &lt;td&gt;ACH-000706&lt;/td&gt;
      &lt;td&gt;PTDSS2&lt;/td&gt;
      &lt;td&gt;1.103052914&lt;/td&gt;
      &lt;td&gt;-1.908065998&lt;/td&gt;
      &lt;td&gt;0.073134&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;NCIH526&lt;/td&gt;
      &lt;td&gt;ACH-000767&lt;/td&gt;
      &lt;td&gt;POU2AF2&lt;/td&gt;
      &lt;td&gt;2.536360826&lt;/td&gt;
      &lt;td&gt;-1.904726778&lt;/td&gt;
      &lt;td&gt;-0.00494779&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;SNU61&lt;/td&gt;
      &lt;td&gt;ACH-000532&lt;/td&gt;
      &lt;td&gt;RB1CC1&lt;/td&gt;
      &lt;td&gt;0.707031795&lt;/td&gt;
      &lt;td&gt;-1.885705392&lt;/td&gt;
      &lt;td&gt;-0.218159067&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;NIHOVCAR3&lt;/td&gt;
      &lt;td&gt;ACH-000001&lt;/td&gt;
      &lt;td&gt;FZR1&lt;/td&gt;
      &lt;td&gt;0.633707063&lt;/td&gt;
      &lt;td&gt;-1.86360478&lt;/td&gt;
      &lt;td&gt;-0.209452144&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;NCIH157DM&lt;/td&gt;
      &lt;td&gt;ACH-000921&lt;/td&gt;
      &lt;td&gt;ARIH2&lt;/td&gt;
      &lt;td&gt;1.084756326&lt;/td&gt;
      &lt;td&gt;-1.852461427&lt;/td&gt;
      &lt;td&gt;0.054123191&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;ZR751&lt;/td&gt;
      &lt;td&gt;ACH-000097&lt;/td&gt;
      &lt;td&gt;ASH1L&lt;/td&gt;
      &lt;td&gt;2.006109006&lt;/td&gt;
      &lt;td&gt;-1.787887575&lt;/td&gt;
      &lt;td&gt;-0.184052968&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;K562&lt;/td&gt;
      &lt;td&gt;ACH-000551&lt;/td&gt;
      &lt;td&gt;MPPE1&lt;/td&gt;
      &lt;td&gt;1.055502649&lt;/td&gt;
      &lt;td&gt;-1.764917658&lt;/td&gt;
      &lt;td&gt;-0.234663579&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;RVH421&lt;/td&gt;
      &lt;td&gt;ACH-000614&lt;/td&gt;
      &lt;td&gt;MAP2K2&lt;/td&gt;
      &lt;td&gt;0.997987991&lt;/td&gt;
      &lt;td&gt;-1.760457331&lt;/td&gt;
      &lt;td&gt;-0.182393707&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;RPMI2650&lt;/td&gt;
      &lt;td&gt;ACH-001385&lt;/td&gt;
      &lt;td&gt;NUTM1&lt;/td&gt;
      &lt;td&gt;0.985643955&lt;/td&gt;
      &lt;td&gt;-1.74866448&lt;/td&gt;
      &lt;td&gt;0.000816&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;SUPT1&lt;/td&gt;
      &lt;td&gt;ACH-000953&lt;/td&gt;
      &lt;td&gt;MAPK14&lt;/td&gt;
      &lt;td&gt;1.189205466&lt;/td&gt;
      &lt;td&gt;-1.742065213&lt;/td&gt;
      &lt;td&gt;-0.174483211&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;UPCISCC090&lt;/td&gt;
      &lt;td&gt;ACH-001227&lt;/td&gt;
      &lt;td&gt;FOXE1&lt;/td&gt;
      &lt;td&gt;17.60251906&lt;/td&gt;
      &lt;td&gt;-1.727963974&lt;/td&gt;
      &lt;td&gt;-0.008935922&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;SF539&lt;/td&gt;
      &lt;td&gt;ACH-000273&lt;/td&gt;
      &lt;td&gt;SPNS1&lt;/td&gt;
      &lt;td&gt;1.223737395&lt;/td&gt;
      &lt;td&gt;-1.727528681&lt;/td&gt;
      &lt;td&gt;-0.283174905&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;FADU&lt;/td&gt;
      &lt;td&gt;ACH-000846&lt;/td&gt;
      &lt;td&gt;C7orf25&lt;/td&gt;
      &lt;td&gt;1.306437402&lt;/td&gt;
      &lt;td&gt;-1.720953978&lt;/td&gt;
      &lt;td&gt;0.038562097&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;SNU739&lt;/td&gt;
      &lt;td&gt;ACH-002523&lt;/td&gt;
      &lt;td&gt;KDELR2&lt;/td&gt;
      &lt;td&gt;1.27255058&lt;/td&gt;
      &lt;td&gt;-1.720720464&lt;/td&gt;
      &lt;td&gt;-0.076599961&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;I then pulled &lt;a href=&quot;https://github.com/armish/vibe-researching/tree/main/depmap-extreme_responders/data/top10deps&quot;&gt;the top 10 dependencies&lt;/a&gt; for each cell line and &lt;a href=&quot;https://github.com/armish/vibe-researching/tree/main/depmap-extreme_responders/data/mutations&quot;&gt;their mutations&lt;/a&gt;; generated &lt;a href=&quot;https://github.com/armish/vibe-researching/tree/main/depmap-extreme_responders/results/prompts&quot;&gt;the prompts&lt;/a&gt;; submitted them programmatically to &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;gpt-5-pro&lt;/code&gt;; collected &lt;a href=&quot;https://github.com/armish/vibe-researching/tree/main/depmap-extreme_responders/results/gpt5-pro-response&quot;&gt;the full responses&lt;/a&gt;; and &lt;a href=&quot;https://github.com/armish/vibe-researching/tree/main/depmap-extreme_responders/results/gpt5-nano-summary&quot;&gt;summarized them&lt;/a&gt; via &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;gpt-5-nano&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Each &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;gpt-5-pro&lt;/code&gt; run took 5–10 minutes and cost about $7. I ran 10 tasks in parallel, so processing 20 extreme responders took ~30 minutes. The total cost, including summarization, was around $150. For what amounts to twenty literature deep dives, this is wild. There is simply no way I could do that manually in half an hour. And yes, there were interesting results. Here are a few highlights:&lt;/p&gt;

&lt;blockquote&gt;

  &lt;p&gt;&lt;strong&gt;SKMES1 is an extreme responder to GSR depletion&lt;/strong&gt; because SK-MES-1 relies almost entirely on the glutathione system because its alternative thioredoxin pathway is crippled by a TXNRD1 splice-acceptor mutation, leaving GSR as the essential antioxidant relay. In other words, the thioredoxin reductase arm is nonfunctional, so loss of GSR eliminates the remaining two-electron reductive capacity. GSR KO would drive GSSG accumulation and widespread protein thiol oxidation, tipping the ER and cytosol toward oxidative damage. The ER is already under heavy proteostasis/oxidative‑folding stress in this cell line, as indicated by dependencies on HYOU1, TMEM258, SLC33A1, and SLC35B4, and a PRKCSH nonsense mutation further strains glycoprotein folding. Mutations that compromise GSH handling (ABCC1 nonsense) and mitochondrial redox buffering (SLC25A24 nonsense) further increase sensitivity, since they reduce GSSG export and mitochondrial ROS buffering. Taken together, these features converge on GSR as the single point of failure for maintaining a reduced GSH pool and a viable redox environment.&lt;/p&gt;

  &lt;p&gt;&lt;strong&gt;EKVX is an extreme responder to PTDSS2 depletion&lt;/strong&gt; because of the paralog synthetic lethality stemming from a truncating PTDSS1 mutation. PTDSS1 and PTDSS2 are the two PS synthases in mammalian cells; either can sustain de novo PS production, but losing both is incompatible with survival. With PTDSS1 function truncated, EKVX must rely on PTDSS2; knocking out PTDSS2 collapses PS biosynthesis and kills the cells. The strong co-dependencies on SELENOI and TMEM258 fit this model: SELENOI supplies PE for PS synthesis, and TMEM258 ties ER function and stress responses to PS flux. This pattern is the classic paralog-based synthetic-lethal logic and matches the data for EKVX. The NSCLC origin of EKVX within the NCI-60 panel is a realistic context for lipid-metabolic rewiring that could reveal such a bottleneck.&lt;/p&gt;

  &lt;p&gt;&lt;strong&gt;UPCISCC090 is an extreme responder to FOXE1 depletion&lt;/strong&gt; because UPCISCC090 has HPV16-driven focal amplifications that include FOXE1, creating a lineage-transcription factor addiction in a keratinocyte/HNSCC context. The HPV integration events flank these amplifications, and the FOXE1 copy number is about sevenfold, which can make FOXE1 essential for survival when perturbed. FOXE1 is a direct GLI2 target in basal keratinocytes and can drive EMT programs through ZEB1, and this particular line also carries a GLI2 missense variant that could tune that axis. Together, the HPV-driven FOXE1 amplification and the Hedgehog/GLI context likely render the cell highly FOXE1-dependent. The broader HPV+ HNSCC wiring—E6–UBE3A, PSMD10/gankyrin, PIM1 amplification, and SOX2 lineage reliance—helps sustain a transcriptional and proteostasis state that becomes brittle when a key amplified TF is removed. In short, a combination of local FOXE1 amplification and keratinocyte lineage signaling best explains the extreme sensitivity to FOXE1 knockout.&lt;/p&gt;

  &lt;p&gt;&lt;strong&gt;RPMI2650 is an extreme responder to NUTM1 depletion&lt;/strong&gt; because RPMI-2650 is driven by a BRD4–NUTM1 fusion, and NUTM1 knockout removes the fusion oncoprotein, collapsing the megadomain-driven chromatin state and transcription program that the fusion creates. Without this program, cells rapidly differentiate toward a squamous lineage and stop growing, which explains the observed growth arrest. This represents a fusion-specific oncogene addiction rather than a general requirement for NUTM1 in somatic cells. The BRD4–NUTM1 ex15:ex2 configuration in RPMI‑2650 is documented, reinforcing its classification as a NUT carcinoma–type dependency. BET inhibitors or p300/CBP inhibitors phenocopy this effect, supporting the link between the fusion-driven transcriptional state and sensitivity to disruption. CERES-corrected CRISPR data further argue that this is a true dependency rather than a copy-number artifact.&lt;/p&gt;

&lt;/blockquote&gt;

&lt;p&gt;There are ~60 more extreme responders (each a one-cell-line, one-gene story) still to go through, and there’s even more to explore: two cell lines sharing the same extreme response, one cell line showing extreme responses across multiple genes, etc. More on that soon. In the meantime, all code, prompts, responses, and summaries are available at&lt;br /&gt;
&lt;a href=&quot;https://github.com/armish/vibe-researching/tree/main/depmap-extreme_responders&quot;&gt;armish/vibe-researching/depmap-extreme_responders&lt;/a&gt;.&lt;/p&gt;
</content>
 </entry>
 
 <entry>
   <title>Announcing {plumber2mcp} and {hgnc.mcp}: of model context protocol and vibe coding</title>
   <link href="https://ergoso.me/ai/vibe/mcp/claude/hgnc/2025/11/21/plumber2mcp-and-hgnc-mcp.html"/>
   <updated>2025-11-21T12:00:00+00:00</updated>
   <id>https://ergoso.me/ai/vibe/mcp/claude/hgnc/2025/11/21/plumber2mcp-and-hgnc-mcp</id>
   <content type="html">&lt;p&gt;On Nov 4, 2025, I received an e-mail from Claude letting me know that I’d been granted $1,000 in credits for the web-based Claude Code environment (&lt;a href=&quot;https://claude.ai/code&quot;&gt;https://claude.ai/code&lt;/a&gt;). The credits were originally set to expire on &lt;del&gt;Nov 18, 2025&lt;/del&gt; (later extended to Nov 23, 2025):&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/img/20251121-1000credits.jpg&quot; alt=&quot;1000 dollar credit from Claude&quot; /&gt;&lt;/p&gt;

&lt;p&gt;Naturally, I decided to make the most of this rare opportunity. I pulled out a few coding projects that had been sitting on the backburner and started vibe-coding as intensely as I could. Surprisingly, despite &lt;strong&gt;furiously&lt;/strong&gt; coding between Nov 4 and Nov 18, I only managed to burn through $89 of the $1,000. I was convinced that such heavy usage would drain the credits quickly, but apparently not:&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/img/20251121-leftovercredits.jpg&quot; alt=&quot;991 dollars left out of 1000 after 2 weeks of intense use&quot; /&gt;&lt;/p&gt;

&lt;h2 id=&quot;plumber2mcp-an-r-package-to-help-convert-a-plumber-api-into-an-mcp-server&quot;&gt;{plumber2mcp}: An R package to help convert a plumber API into an MCP server&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://github.com/armish/plumber2mcp/&quot;&gt;Code&lt;/a&gt; / &lt;a href=&quot;https://arman.aksoy.org/plumber2mcp/&quot;&gt;Documents&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;One of my first experiments was inspired by &lt;a href=&quot;https://github.com/tadata-org/fastapi_mcp&quot;&gt;FastAPI-MCP&lt;/a&gt;. I loved the idea of taking an existing &lt;a href=&quot;https://github.com/fastapi/fastapi&quot;&gt;FastAPI&lt;/a&gt; application and essentially “turning it into” an MCP server. I wanted to do the same for R: take a &lt;a href=&quot;https://www.rplumber.io/&quot;&gt;plumber&lt;/a&gt; API and expose its endpoints as MCP utilities. It also seemed like a good excuse to get some hands-on exposure to MCP, because at the time, the protocol felt confusing and abstract.&lt;/p&gt;

&lt;p&gt;I’m a long-time fan of &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;plumber&lt;/code&gt; and have used it extensively to build internally facing APIs. So back in July, when I first got access to Claude Code, I started with a simple prompt: “Build an R package that takes a plumber object and exposes its endpoints via MCP.” To make things easier for Claude—since R isn’t its strongest suit—I prepared the package skeleton with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;usethis&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Within a few days of intermittent vibing, I had a minimally functioning package. Testing the MCP side was more painful than expected, especially the requirement of wrapping HTTP endpoints with a &lt;a href=&quot;https://github.com/armish/plumber2mcp/blob/main/inst/examples/stdio-wrapper.py&quot;&gt;script&lt;/a&gt; to support STDIO transport mode. It wasn’t shocking given how early MCP tooling still is, but I did hope for something a bit more plug-and-play. Hopefully the package documentation will make the experience smoother for others who want to build MCP tools.&lt;/p&gt;

&lt;p&gt;Here’s what a typical plumber → MCP conversion looks like using &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;plumber2mcp&lt;/code&gt;:&lt;/p&gt;

&lt;div class=&quot;language-r highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;n&quot;&gt;library&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;plumber&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;library&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;plumber2mcp&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;

&lt;/span&gt;&lt;span class=&quot;c1&quot;&gt;# Read API endpoint definitions, add MCP endpoints, and then run it&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;pr&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;plumber.R&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;%&amp;gt;%&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;pr_mcp&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;transport&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;http&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;%&amp;gt;%&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;pr_run&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;port&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;m&quot;&gt;8000&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;That first project taught me a lot about MCP. It was fun to watch trivial functions show up inside the MCP inspector and behave as expected. But I knew that for a serious application, I’d eventually need to revisit and refine the package.&lt;/p&gt;

&lt;h2 id=&quot;hgncmcp-a-combined-api-and-mcp-server-to-ease-integration-of-hgnc-resources-with-llms&quot;&gt;{hgnc.mcp}: A combined API and MCP server to ease integration of HGNC resources with LLMs&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://github.com/armish/hgnc.mcp/&quot;&gt;Code&lt;/a&gt; / &lt;a href=&quot;https://arman.aksoy.org/hgnc.mcp/&quot;&gt;Documents&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;As a bioinformatics researcher, I rely heavily on the &lt;a href=&quot;https://www.genenames.org/&quot;&gt;HUGO Gene Nomenclature Committee (HGNC)&lt;/a&gt;. LLMs—ChatGPT, Claude, and others—are generally knowledgeable about human genes, but they often struggle with alias resolution, gene families, or external IDs. This becomes a real problem during deep research tasks, where a model can get stuck on an outdated synonym or miss information due to a recent name change.&lt;/p&gt;

&lt;p&gt;Two years ago, I attempted to address this by building a custom GPT called &lt;a href=&quot;https://chatgpt.com/g/g-E7GrLSwnv-hugo-bioinformatics-helper&quot;&gt;Hugo - Bioinformatics helper&lt;/a&gt;. It worked well for normalization and lookup tasks, but it was locked inside ChatGPT and couldn’t be integrated into other LLMs.&lt;/p&gt;

&lt;div style=&quot;width: 50%; margin: 0 auto;&quot;&gt;&lt;blockquote class=&quot;twitter-tweet&quot;&gt;&lt;p lang=&quot;en&quot; dir=&quot;ltr&quot;&gt;Just got GPTs enabled for my account and, of course, the first GPT I tried to build is an &lt;a href=&quot;https://twitter.com/genenames?ref_src=twsrc%5Etfw&quot;&gt;@genenames&lt;/a&gt;-savvy assistant.&lt;br /&gt;&lt;br /&gt;Super excited to finally have a companion who can help me with one-off mundane ID mapping/look-up tasks &lt;a href=&quot;https://t.co/LZYb9ffeGT&quot;&gt;pic.twitter.com/LZYb9ffeGT&lt;/a&gt;&lt;/p&gt;&amp;mdash; B. Arman Aksoy (@armish) &lt;a href=&quot;https://twitter.com/armish/status/1722629935096103069?ref_src=twsrc%5Etfw&quot;&gt;November 9, 2023&lt;/a&gt;&lt;/blockquote&gt; &lt;script async=&quot;&quot; src=&quot;https://platform.twitter.com/widgets.js&quot; charset=&quot;utf-8&quot;&gt;&lt;/script&gt;&lt;/div&gt;

&lt;p&gt;With the free Claude credits acting as a catalyst, I wanted to revisit the idea—this time as an MCP server. I created an empty repository, opened a conversation with Claude, and outlined the plan:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;build an R package depending on &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;plumber2mcp&lt;/code&gt;,&lt;/li&gt;
  &lt;li&gt;expose HGNC resources as MCP tools,&lt;/li&gt;
  &lt;li&gt;implement alias mapping, external ID lookups, gene family retrieval, etc.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I asked Claude not to write code immediately but instead to design the plan as a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;TODO.md&lt;/code&gt;. For each item in that TODO list, I opened a fresh conversation, let it generate a PR, and merged when things looked good.&lt;/p&gt;

&lt;p&gt;Working on an R project inside Claude Code Web has limitations. The environment includes basic Python but no R installation, which means most of the R code was produced completely “blind.” Actual testing mostly happened in GitHub Actions, so I ended up bouncing between action logs and Claude Code Web—definitely not the smoothest workflow. On the command line, Claude Code can leverage your local environment and iterate much faster.&lt;/p&gt;

&lt;p&gt;But still, I was &lt;strong&gt;impressed&lt;/strong&gt; with how much blind coding actually worked. Most tests passed on the first try, and when something failed, pasting the error into Claude was usually enough to fix it on the spot.&lt;/p&gt;

&lt;p&gt;This wasn’t a small project either. It required:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;creating a proper R package,&lt;/li&gt;
  &lt;li&gt;learning how to work with HGNC’s TSV files,&lt;/li&gt;
  &lt;li&gt;providing a suite of functions,&lt;/li&gt;
  &lt;li&gt;decorating them with &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;plumber&lt;/code&gt; annotations,&lt;/li&gt;
  &lt;li&gt;exposing them as MCP tools/resources,&lt;/li&gt;
  &lt;li&gt;and testing everything end-to-end.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Ten days of intense vibe-coding later, I had spent only $81 and ended up with a solid R package that did exactly what I needed.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/img/20251121-claudecodeweb.jpg&quot; alt=&quot;Claude Code Web in action&quot; /&gt;&lt;/p&gt;

&lt;p&gt;When it was time to test integration with LLMs, things got complicated again. I wanted to use Claude Desktop, but expecting normal users to install R just to access HGNC data wasn’t realistic. So I decided to ship everything as a Docker image and use the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;STDIO&lt;/code&gt; transport. This required switching to Claude Code CLI because waiting for GitHub Actions to build and push images before locally testing was too slow.&lt;/p&gt;

&lt;p&gt;Once the Docker workflow was set up, I discovered that my earlier &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;plumber2mcp&lt;/code&gt; implementation was based on an outdated MCP specification. Updating it to the newer version wasn’t too bad, though, and before long, I had everything running smoothly in Claude Desktop:&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/img/20251121-hgncmcpmenu.jpg&quot; alt=&quot;HGNC MCP submenu under Claude Desktop tool settings&quot; /&gt;&lt;/p&gt;

&lt;p&gt;Here’s an example showing how the HGNC MCP server resolves a mix of aliases without needing explicit instructions—and without worrying about outdated knowledgebases:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;What are the official gene symbols and entrez gene ids for PD-L1, PD-1, TCF-1, OX40, CD3?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;img src=&quot;/img/20251121-hgncmcpdemo.jpg&quot; alt=&quot;HGNC MCP server in action within Claude Desktop&quot; /&gt;&lt;/p&gt;

&lt;p&gt;If you’re a Claude Desktop user and want to try it yourself, first pull the latest Docker image:&lt;/p&gt;

&lt;div class=&quot;language-bash highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;nv&quot;&gt;$ &lt;/span&gt;docker pull ghcr.io/armish/hgnc.mcp:latest
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Then update your Claude configuration (on macOS this lives at &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;~/Library/Application\ Support/Claude/claude_desktop_config.json&lt;/code&gt;):&lt;/p&gt;

&lt;div class=&quot;language-json highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;&quot;mcpServers&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
    &lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;&quot;hgnc&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
      &lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;&quot;command&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;docker&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
      &lt;/span&gt;&lt;span class=&quot;nl&quot;&gt;&quot;args&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;run&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;--rm&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;-i&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;-v&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;hgnc-cache:/home/hgnc/.cache/hgnc&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;ghcr.io/armish/hgnc.mcp:latest&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;--stdio&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
    &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;After restarting Claude Desktop, you should see &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;hgnc&lt;/code&gt; under your tools menu, where you can toggle individual tools on or off.&lt;/p&gt;

&lt;h2 id=&quot;final-thoughts&quot;&gt;Final thoughts&lt;/h2&gt;

&lt;p&gt;What did I learn from this extended vibe-coding adventure?&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;It has &lt;strong&gt;never&lt;/strong&gt; been this easy to go from a rough idea to a functional solution. If the credits were coming out of my own pocket, it might feel different—but for ~$90, the experience was absolutely worth it.&lt;/li&gt;
  &lt;li&gt;Building an MCP solution is not trivial. It spans multiple layers (R package → API → MCP → STDIO wrapping → Dockerization → LLM configuration) and requires constant context-switching.&lt;/li&gt;
  &lt;li&gt;Getting an MCP server to “work” is only half the story: a couple of queries can blow through the conversation’s context window. Optimizing MCP responses for token efficiency is its own challenge.&lt;/li&gt;
  &lt;li&gt;Claude Code Web is great for lightweight work, but blind-coding R packages gets clumsy fast. Claude Code CLI is significantly smoother because it can run things locally and iterate quickly. If only my free credits applied there too.&lt;/li&gt;
  &lt;li&gt;The limbo periods while LLMs are thinking are real. For quick tasks, the pauses are manageable. But during long vibe-coding sessions, the waiting becomes mentally draining. Session blocks help, but it’s still easy to burn out.&lt;/li&gt;
  &lt;li&gt;Publishing an R package on CRAN requires patience. I submitted an early version of &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;plumber2mcp&lt;/code&gt; almost two weeks ago and still haven’t heard back. Until then, both packages are available on GitHub and can be installed via &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;devtools&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;remotes&lt;/code&gt;, etc.&lt;/li&gt;
&lt;/ul&gt;
</content>
 </entry>
 
 <entry>
   <title>Transcript: The Repertoire Room Episode 3 - TCR-sequencing</title>
   <link href="https://ergoso.me/immunewatch/detect/tcr/translational/computational/bioinformatics/sequencing/2025/10/28/immunewatch-repertoire-room-episode-3.html"/>
   <updated>2025-10-28T00:00:01+00:00</updated>
   <id>https://ergoso.me/immunewatch/detect/tcr/translational/computational/bioinformatics/sequencing/2025/10/28/immunewatch-repertoire-room-episode-3</id>
   <content type="html">&lt;div style=&quot;position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden; max-width: 100%;&quot;&gt;
  &lt;iframe src=&quot;https://www.youtube.com/embed/J41q0fUqeVo&quot; frameborder=&quot;0&quot; allow=&quot;accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture&quot; allowfullscreen=&quot;&quot; style=&quot;position: absolute; top: 0; left: 0; width: 100%; height: 100%;&quot;&gt;
  &lt;/iframe&gt;
&lt;/div&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;interview-transcript&quot;&gt;Interview Transcript&lt;/h2&gt;
&lt;p&gt;&lt;em&gt;Edited for clarity and readability&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sander Wuyts:&lt;/strong&gt;&lt;br /&gt;
Hi everyone. A couple of years ago, I became completely obsessed with T cell repertoires—and I know I’m not the only one. That’s why I started this podcast. Welcome to &lt;strong&gt;The Repertoire Room&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;My name is Sander, and I’m the co-founder and CEO of ImmuneWatch, an immunoinformatics company specializing in T cell receptor sequencing analysis. In this show, I aim to have insightful conversations with people who share this passion and go deep into T cell sequencing technologies and analysis.&lt;/p&gt;

&lt;p&gt;Today, I’m very excited to welcome &lt;strong&gt;Arman Aksoy&lt;/strong&gt;, Director of Computational Biology at Obsidian Therapeutics. Welcome, Arman.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Arman Aksoy:&lt;/strong&gt;&lt;br /&gt;
Thank you so much for having me, Sander.&lt;/p&gt;

&lt;hr /&gt;

&lt;p&gt;&lt;strong&gt;Sander Wuyts:&lt;/strong&gt;&lt;br /&gt;
I’m excited to have you here. Let’s start by introducing you to everyone. What is your current role?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Arman Aksoy:&lt;/strong&gt;&lt;br /&gt;
I’m the Director of Computational Biology at Obsidian Therapeutics. My role involves coordinating data generation and analysis efforts across early discovery, preclinical, translational, and clinical functions.&lt;/p&gt;

&lt;p&gt;As a startup, we’re a very lean team. My department officially consists of one person—myself—but I work very closely with at least ten others at the company who contribute computationally to the biology work.&lt;/p&gt;

&lt;hr /&gt;

&lt;p&gt;&lt;strong&gt;Sander Wuyts:&lt;/strong&gt;&lt;br /&gt;
I love the “one-person department” description. That feels very fitting for a startup.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Arman Aksoy:&lt;/strong&gt;&lt;br /&gt;
Absolutely. Times are changing. There’s a new wave of scientists entering both academia and industry with strong computational backgrounds, and I think the computational biology landscape is constantly evolving.&lt;/p&gt;

&lt;hr /&gt;

&lt;p&gt;&lt;strong&gt;Sander Wuyts:&lt;/strong&gt;&lt;br /&gt;
Many people enter computational biology through different paths. What’s your origin story?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Arman Aksoy:&lt;/strong&gt;&lt;br /&gt;
It definitely wasn’t a straight path. I’m originally from Turkey, where I lived and studied until the end of college. At the time, there were no formal bioinformatics or computational biology programs. I double-majored in cell biology and mathematics, which helped shape me into a hybrid scientist.&lt;/p&gt;

&lt;p&gt;I worked in a veterinary lab focused on neurodegenerative diseases, then joined Memorial Sloan Kettering Cancer Center for graduate school. My training coincided with the rise of The Cancer Genome Atlas, and I became deeply immersed in genomics and translational medicine.&lt;/p&gt;

&lt;p&gt;Immunology initially intimidated me, but during my postdoctoral work at the Icahn School of Medicine—at the height of checkpoint blockade therapies—I found myself drawn into immuno-oncology. Our lab evolved from purely computational into a hybrid lab, and I returned to wet-lab work focusing on T cell biology and adaptive cell therapies.&lt;/p&gt;

&lt;p&gt;From there, I joined Agenus, then MOMA Therapeutics as the founding computational biologist, and eventually Obsidian Therapeutics, where I work today.&lt;/p&gt;

&lt;hr /&gt;

&lt;p&gt;&lt;strong&gt;Sander Wuyts:&lt;/strong&gt;&lt;br /&gt;
Do you think being a hybrid scientist helped you in data analysis?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Arman Aksoy:&lt;/strong&gt;&lt;br /&gt;
Absolutely. Understanding how data is generated—protocols, sample preparation, and failure points—makes analysis far more effective. Wet-lab experience helps you recognize sources of error and design better experiments.&lt;/p&gt;

&lt;hr /&gt;

&lt;p&gt;&lt;strong&gt;Sander Wuyts:&lt;/strong&gt;&lt;br /&gt;
What does Obsidian Therapeutics do today?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Arman Aksoy:&lt;/strong&gt;&lt;br /&gt;
We’re a cell and gene therapy company. Our lead program, &lt;strong&gt;OBX-115&lt;/strong&gt;, is an investigational tumor-infiltrating lymphocyte (TIL) therapy.&lt;/p&gt;

&lt;p&gt;What makes it unique is that the TILs are engineered to express membrane-bound IL-15, removing the need for IL-2 support and reducing toxicity. We can also regulate IL-15 expression using a small molecule. OVX115 is currently in Phase II clinical trials.&lt;/p&gt;

&lt;hr /&gt;

&lt;p&gt;&lt;strong&gt;Sander Wuyts:&lt;/strong&gt;&lt;br /&gt;
How do T cell repertoires factor into your work?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Arman Aksoy:&lt;/strong&gt;&lt;br /&gt;
TIL therapy gives us an incredible opportunity to deeply profile polyclonal T cell repertoires. We perform routine TCR sequencing on infusion products and track clonotypes over time in blood and tissue samples.&lt;/p&gt;

&lt;p&gt;This allows us to study persistence, expansion, depletion, and potential biomarkers of response. These data are central to our translational efforts.&lt;/p&gt;

&lt;hr /&gt;

&lt;p&gt;&lt;strong&gt;Sander Wuyts:&lt;/strong&gt;&lt;br /&gt;
What were the biggest challenges in implementing TCR sequencing?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Arman Aksoy:&lt;/strong&gt;&lt;br /&gt;
Feasibility studies are critical—you can’t treat TCR sequencing as an afterthought. You need to test platforms, sequencing depth, vendors, and computational tools upfront.&lt;/p&gt;

&lt;p&gt;Building a large internal dataset is essential for interpretation. TCR data are complex, and summary metrics like diversity and clonality can be misleading without proper context.&lt;/p&gt;

&lt;p&gt;Annotation remains one of the hardest challenges. Mapping TCRs to antigens is still evolving, so we rely on validated external tools rather than constantly rebuilding internal solutions.&lt;/p&gt;

&lt;hr /&gt;

&lt;p&gt;&lt;strong&gt;Sander Wuyts:&lt;/strong&gt;&lt;br /&gt;
Did all this effort make a difference?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Arman Aksoy:&lt;/strong&gt;&lt;br /&gt;
Absolutely. It allowed us to deeply understand our lead product across preclinical and clinical settings. TCR sequencing helps us troubleshoot, optimize, and interpret product behavior in ways that wouldn’t otherwise be possible.&lt;/p&gt;

&lt;hr /&gt;

&lt;p&gt;&lt;strong&gt;Sander Wuyts:&lt;/strong&gt;&lt;br /&gt;
How do you see the field evolving?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Arman Aksoy:&lt;/strong&gt;&lt;br /&gt;
TCR-to-epitope datasets are still sparse, but even a single confident annotation can anchor interpretation. I expect improved experimental platforms and integration with single-cell transcriptomics to significantly advance the field.&lt;/p&gt;

&lt;p&gt;Ultimately, I see a future where we can sequence TCRs and meaningfully understand their targets and therapeutic relevance.&lt;/p&gt;

&lt;hr /&gt;

&lt;p&gt;&lt;strong&gt;Sander Wuyts:&lt;/strong&gt;&lt;br /&gt;
Final question: if you weren’t a computational biologist, what would you be doing?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Arman Aksoy:&lt;/strong&gt;&lt;br /&gt;
I’d be in the lab. I miss experimental work and being a true hybrid scientist.&lt;/p&gt;

&lt;hr /&gt;

&lt;p&gt;&lt;strong&gt;Sander Wuyts:&lt;/strong&gt;&lt;br /&gt;
Where can people learn more about you and Obsidian?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Arman Aksoy:&lt;/strong&gt;&lt;br /&gt;
Our website, &lt;strong&gt;https://obsidiantx.com&lt;/strong&gt;, is the best place to learn about Obsidian. I’m also active on LinkedIn, BlueSky, GitHub, X, and Hugging Face—just search my name.&lt;/p&gt;
</content>
 </entry>
 
 <entry>
   <title>The Deceptively Complex World of Turkish Diacritics: A Neural Network Journey</title>
   <link href="https://ergoso.me/turkish/neural/network/github/diacritics/macos/app/swift/python/2025/09/17/turkish-diacritic-restoration.html"/>
   <updated>2025-09-17T18:00:01+00:00</updated>
   <id>https://ergoso.me/turkish/neural/network/github/diacritics/macos/app/swift/python/2025/09/17/turkish-diacritic-restoration</id>
   <content type="html">&lt;p&gt;&lt;em&gt;How adding a few dots and curves to letters became a 3.7‑million‑parameter adventure&lt;/em&gt;&lt;/p&gt;

&lt;hr /&gt;

&lt;h2 id=&quot;the-simple-problem&quot;&gt;The “Simple” Problem&lt;/h2&gt;

&lt;p&gt;Picture this: you’re reading Turkish text where someone forgot to add the special characters. “Turkiye” instead of “Türkiye”, “ogrenci” instead of “öğrenci.” Seems simple enough, right? Just add back the missing dots and curves – how hard could it be?&lt;/p&gt;

&lt;p&gt;Well, buckle up. What looks like a search‑and‑replace task quickly becomes a journey through morphology, orthography, GPU quirks, and model design.&lt;/p&gt;

&lt;h2 id=&quot;what-are-diacritics-anyway&quot;&gt;What Are Diacritics, Anyway?&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Diacritics&lt;/em&gt; are the marks you see on letters that change their sound or meaning - like &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;é&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ç&lt;/code&gt;, or &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ö&lt;/code&gt;. In Turkish, six Latin letters have diacritic counterparts: &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;c↔ç&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;g↔ğ&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;i↔ı&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;o↔ö&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;s↔ş&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;u↔ü&lt;/code&gt;. Missing a diacritic is not just a cosmetic issue: it can turn a word into a different word (“yasam” → “yaşam”), or a non‑word into a valid word (“semşiye” → “şemsiye”). That means you need &lt;strong&gt;context&lt;/strong&gt; to decide what the right letter should be.&lt;/p&gt;

&lt;h2 id=&quot;origin-story-from-a-mac-app-idea-to-a-deep-dive&quot;&gt;Origin Story: From a Mac App Idea to a Deep Dive&lt;/h2&gt;

&lt;p&gt;This project didn’t start as a research exercise. I originally just wanted to build a &lt;strong&gt;macOS app&lt;/strong&gt; that restores diacritics naively - a little utility I could run on local text. While searching for inspiration, I found &lt;a href=&quot;https://ileriseviye.wordpress.com/tag/turkish-deasciifier/&quot;&gt;Emre Sevinç’s blog post on Turkish deasciification&lt;/a&gt;, which pointed to &lt;a href=&quot;https://github.com/emres/turkish-deasciifier&quot;&gt;his own Python implementation of the classic pattern‑based approach&lt;/a&gt;, and also gathered the &lt;strong&gt;historical context&lt;/strong&gt; of the problem. That post also name‑checked &lt;strong&gt;neural networks&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Cue the curiosity. If rule‑based tools could already do well, what could a modern neural network approach do on a standard laptop? After vibe coding a macOS application in Swift (&lt;a href=&quot;https://github.com/armish/TurkishDeasciifier&quot;&gt;https://github.com/armish/TurkishDeasciifier&lt;/a&gt;), I decided to vibe‑code my way into neural networks and see how far I could push it.&lt;/p&gt;

&lt;p&gt;For the impatient, here is the repository that has all the relevant materials (code, models, more details): &lt;a href=&quot;https://github.com/armish/nokta-ai&quot;&gt;https://github.com/armish/nokta-ai&lt;/a&gt;.&lt;/p&gt;

&lt;h2 id=&quot;a-15year-head-start-rulebased-deasciifiers&quot;&gt;A 15‑Year Head Start: Rule‑Based Deasciifiers&lt;/h2&gt;

&lt;p&gt;Long before GPUs were in the picture, researchers built &lt;strong&gt;pattern‑based deasciifiers&lt;/strong&gt; that worked surprisingly well on CPUs: for example, &lt;strong&gt;Deniz Yüret’s Emacs Lisp Turkish mode&lt;/strong&gt; (one of the earliest practical tools). These systems were &lt;strong&gt;deterministic, blazing fast, and accurate (~97%)&lt;/strong&gt; for everyday use. So why revisit the problem? Because by the 2020s, the landscape shifted: M‑series laptops with a GPU‑like Metal backend, cheap on‑demand cloud GPUs, and AI copilots that accelerate coding. What used to take months could plausibly happen over a weekend.&lt;/p&gt;

&lt;h2 id=&quot;stepwise-development-wrestling-with-neural-nets&quot;&gt;Stepwise Development: Wrestling With Neural Nets&lt;/h2&gt;

&lt;p&gt;My path to a strong model was iterative, and each step fixed a very specific pain point:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;
    &lt;p&gt;&lt;strong&gt;Naïve BiLSTM&lt;/strong&gt;: A basic character‑level recurrent model. Result: underwhelming accuracy, especially on ambiguous words.&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;&lt;strong&gt;Mask non‑diacritic characters&lt;/strong&gt;: Why waste capacity predicting diacritics for letters like &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;m&lt;/code&gt; or &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;t&lt;/code&gt;? Masking &lt;strong&gt;reduced the output space&lt;/strong&gt; and &lt;strong&gt;cut false positives&lt;/strong&gt;. This was a crucial simplifying step &lt;em&gt;before&lt;/em&gt; playing with attention.&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;&lt;strong&gt;Add an attention layer&lt;/strong&gt;: Attention let the model focus on relevant context (prefixes/suffixes, neighboring words). Helpful, but not a silver bullet.&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;&lt;strong&gt;Balanced sampling + weighted loss&lt;/strong&gt;: Some diacritics are rare. Without balancing and a &lt;strong&gt;weighted BCE loss&lt;/strong&gt;, the model under‑predicted them. Oversampling + loss weights produced a &lt;strong&gt;big jump&lt;/strong&gt; in accuracy.&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;&lt;strong&gt;Reduce to 6 characters (lowercase normalization)&lt;/strong&gt;: I normalized the corpus to lowercase and trained the model to decide &lt;strong&gt;only&lt;/strong&gt; among the six diacritic pairs. After inference, I &lt;strong&gt;restored case&lt;/strong&gt; using explicit Turkish rules (including the dotted/dotless &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;i&lt;/code&gt;). This &lt;strong&gt;massively simplified&lt;/strong&gt; learning.&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;&lt;strong&gt;Special handling for the dotted and dotless i&lt;/strong&gt;: English uppercases &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;i&lt;/code&gt; to &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;I&lt;/code&gt;. Turkish does something different: lowercase &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;i&lt;/code&gt; becomes uppercase &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;İ&lt;/code&gt; (dotted I) and lowercase &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ı&lt;/code&gt; (dotless i) becomes uppercase &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;I&lt;/code&gt;. That means &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;istanbul&lt;/code&gt; uppercased correctly is &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;İSTANBUL&lt;/code&gt;, not &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ISTANBUL&lt;/code&gt;. Words like &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ışık&lt;/code&gt; (light) and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;işik&lt;/code&gt; are different words. I normalized everything to lowercase for training, then restored case in post-processing with explicit rules for &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;i/İ/ı/I&lt;/code&gt;. This avoided teaching the network case rules and eliminated one of the biggest error sources.&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;&lt;strong&gt;Scale on a GPU&lt;/strong&gt;: With the fundamentals fixed, increasing the &lt;strong&gt;context window&lt;/strong&gt;, &lt;strong&gt;hidden size&lt;/strong&gt;, and especially the &lt;strong&gt;dataset size&lt;/strong&gt; on an &lt;strong&gt;A100&lt;/strong&gt; unlocked near‑SOTA accuracy.&lt;/p&gt;
  &lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Each tweak felt like peeling away a layer of complexity until things finally clicked.&lt;/p&gt;

&lt;h2 id=&quot;the-architecture-that-finally-worked-highlevel&quot;&gt;The Architecture That Finally Worked (High‑Level)&lt;/h2&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Backbone:&lt;/strong&gt; 2‑layer &lt;strong&gt;BiLSTM&lt;/strong&gt; with a modest &lt;strong&gt;hidden state&lt;/strong&gt; per direction.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Attention:&lt;/strong&gt; lightweight multi‑head self‑attention to let the model “peek” farther than the recurrent state.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Context window:&lt;/strong&gt; up to &lt;strong&gt;96&lt;/strong&gt; characters in the A100 run – enough for word boundaries and morphology.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Output heads:&lt;/strong&gt; six binary classifiers (one per diacritic pair), gated by the &lt;strong&gt;mask&lt;/strong&gt; so only eligible characters are considered.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Training:&lt;/strong&gt; &lt;strong&gt;balanced sampling&lt;/strong&gt;, &lt;strong&gt;weighted loss&lt;/strong&gt;, lowercase normalization, explicit &lt;strong&gt;post‑processing&lt;/strong&gt; for Turkish casing rules (especially &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;i/İ/ı/I&lt;/code&gt;).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This blend kept the model &lt;strong&gt;small enough to ship&lt;/strong&gt; (4.4–16 MB) while retaining the context it needs to disambiguate tricky cases.&lt;/p&gt;

&lt;h2 id=&quot;datasets--test-files&quot;&gt;Datasets &amp;amp; Test Files&lt;/h2&gt;

&lt;p&gt;To keep things honest across domains, I evaluated on three very different test sets:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;aysnrgenc_turkishdeasciifier_test.txt&lt;/code&gt;&lt;/strong&gt; - a subset derived from &lt;a href=&quot;https://github.com/aysnrgenc/turkishdeasciifier&quot;&gt;Aysenur Genç’s implementation&lt;/a&gt;.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;llm_random_test.txt&lt;/code&gt;&lt;/strong&gt; - random Turkish text generated by ChatGPT (noisy, synthetic).&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;vikipedi_test.txt&lt;/code&gt;&lt;/strong&gt; - long, stitched Turkish Wikipedia articles (cleaner, formal, long‑context).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For evaluation, I focused on two &lt;strong&gt;meaningful&lt;/strong&gt; metrics:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Diacritic‑specific accuracy&lt;/strong&gt; - &lt;em&gt;did it restore where it needed to?&lt;/em&gt;&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Overall word accuracy&lt;/strong&gt; - &lt;em&gt;did it avoid changing words that were already correct?&lt;/em&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2 id=&quot;laptop-vs-gpu-training-realities&quot;&gt;Laptop vs. GPU: Training Realities&lt;/h2&gt;

&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Environment&lt;/th&gt;
      &lt;th style=&quot;text-align: right&quot;&gt;Context window&lt;/th&gt;
      &lt;th style=&quot;text-align: right&quot;&gt;Hidden size&lt;/th&gt;
      &lt;th style=&quot;text-align: right&quot;&gt;Batch size&lt;/th&gt;
      &lt;th style=&quot;text-align: right&quot;&gt;Train set&lt;/th&gt;
      &lt;th style=&quot;text-align: right&quot;&gt;Epochs&lt;/th&gt;
      &lt;th&gt;Runtime&lt;/th&gt;
      &lt;th style=&quot;text-align: right&quot;&gt;Overall word accuracy&lt;/th&gt;
      &lt;th style=&quot;text-align: right&quot;&gt;Diacritic-specific accuracy&lt;/th&gt;
      &lt;th style=&quot;text-align: right&quot;&gt;Model size&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;strong&gt;M1 Pro (overnight prototype)&lt;/strong&gt;&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;20&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;128&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;16&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;10,000&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;50&lt;/td&gt;
      &lt;td&gt;~10 h&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;83–93%&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;82–92%&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;4.4 MB&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;strong&gt;A100 GPU (small config)&lt;/strong&gt;&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;20&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;128&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;32&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;100,000&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;50&lt;/td&gt;
      &lt;td&gt;&amp;lt;10 h&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;~99%&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;~99%&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;4.4 MB&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;&lt;strong&gt;A100 GPU (scaled config)&lt;/strong&gt;&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;96&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;256&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;128&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;100,000&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;50&lt;/td&gt;
      &lt;td&gt;&amp;lt;10 h&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;~99–99.5%&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;~99–99.6%&lt;/td&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;16 MB&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;&lt;strong&gt;Why this matters:&lt;/strong&gt; Data often beats parameters - simply training the small architecture on 100k sentences closed most of the gap. Doing the same 10x dataset on an M1 Pro would take more than a week, so this result is only practical because of the A100’s throughput.&lt;/p&gt;

&lt;h2 id=&quot;crossdevice-plot-twist-same-weights-different-results&quot;&gt;Cross‑Device Plot Twist: Same Weights, Different Results&lt;/h2&gt;

&lt;p&gt;When I moved an A100‑trained model to my MacBook (MPS backend), its accuracy dropped. Same weights - different behavior. Why? Apparently,&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Precision modes&lt;/strong&gt; differ (A100 often uses &lt;strong&gt;TF32&lt;/strong&gt; by default; MPS uses &lt;strong&gt;FP32&lt;/strong&gt;).&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Backend kernels&lt;/strong&gt; differ (cuDNN vs. Metal), changing reduction order and numerical drift.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;MPS CPU fallbacks&lt;/strong&gt; for unsupported ops can subtly change numerics.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So, addressing this is still on my TODO list but while before I do that, both the MPS- and CUDA-compatible models are available under &lt;a href=&quot;https://github.com/armish/nokta-ai/releases/tag/v0.1.0&quot;&gt;the v0.1.0 release&lt;/a&gt;.&lt;/p&gt;

&lt;h2 id=&quot;lessons-learned&quot;&gt;Lessons Learned&lt;/h2&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;strong&gt;Simple problems aren’t simple.&lt;/strong&gt; Language has corner cases; Turkish has &lt;em&gt;delightful&lt;/em&gt; ones.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Masking and data balance matter more than fancy layers.&lt;/strong&gt; You can’t out‑architect skewed data.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Constraints breed innovation.&lt;/strong&gt; Lowercasing + six‑character focus + explicit case restoration beat my 12‑head uppercase/lowercase attempt by a mile.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Hardware matters.&lt;/strong&gt; GPUs don’t just speed things up; they expose bugs and numerical differences.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Data is king.&lt;/strong&gt; Scaling examples produced bigger gains than scaling parameters.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Cross‑platform reproducibility is tricky.&lt;/strong&gt; CUDA and MPS won’t perfectly agree.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2 id=&quot;final-thought-fun-over-necessity&quot;&gt;Final Thought: Fun Over Necessity&lt;/h2&gt;

&lt;p&gt;Rule‑based deasciifiers solved this well enough for most use cases &lt;strong&gt;years&lt;/strong&gt; ago - fast, deterministic, CPU‑only. Neural nets can push accuracy closer to perfection, but they demand GPUs, data pipelines, and overnight runs.&lt;/p&gt;

&lt;p&gt;Was any of this strictly necessary? &lt;strong&gt;No.&lt;/strong&gt; Was it a blast? &lt;strong&gt;Absolutely.&lt;/strong&gt; 🚀&lt;/p&gt;

&lt;p&gt;The real joy wasn’t “beating” the old tools; it was the journey - discovering quirks of Turkish orthography, debugging across hardware, and watching a model learn that &lt;em&gt;“kasap”&lt;/em&gt; stays as &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;s&lt;/code&gt; while &lt;em&gt;“yaşam”&lt;/em&gt; becomes &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ş&lt;/code&gt;.&lt;/p&gt;
</content>
 </entry>
 
 <entry>
   <title>evSrc: Evolutionary couplings between files reveal poor design choices in software architecture</title>
   <link href="https://ergoso.me/computer/science/github/software/evolutionary/couplings/2014/12/10/evsrc-evolutionary-couplings-reveal-poor-software-design.html"/>
   <updated>2014-12-10T13:00:01+00:00</updated>
   <id>https://ergoso.me/computer/science/github/software/evolutionary/couplings/2014/12/10/evsrc-evolutionary-couplings-reveal-poor-software-design</id>
   <content type="html">&lt;p&gt;Given a software code repository together with all commit history, can you spot potentially problematic parts of the software?
This is the question I have been thinking about when I have some extra time left from other projects.
As a pet-project, I recently created a &lt;em&gt;proof-of-concept&lt;/em&gt; pipeline to show that the answer to this question seems to be &lt;strong&gt;yes&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;I was thinking about writing up a small application paper for this project,
but I am really terrible at reading papers from the Computer Science field, let alone writing them.
I also realize that I barely have the time to pursue this project any further,
so instead of turning it into a dead project, I decided to write a blog post about the current state of it.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;evSrc&lt;/em&gt;, or EVolutionary SouRCe, project came into being while I was talking to &lt;a href=&quot;http://cbio.mskcc.org/directory/richard-stein/index.html&quot;&gt;Richard&lt;/a&gt; about the details of maximum entropy-based inference of pairwise interactions from large-scale data sets.
In case you missed it, 
there have been really exciting developments in the Structural Biology field,
where it has been shown that using this method and taking advantage of publicly available sequencing information,
you can fold proteins and you can even find structures of complexes when you have enough sequences for protein(s).
You can learn more about these projects from the &lt;a href=&quot;http://www.evfold.org&quot;&gt;EVFold website&lt;/a&gt;, but the main pipeline looks something like this:&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;http://www.evfold.org&quot;&gt;&lt;img src=&quot;/img/evsrc-evfold.png&quot; alt=&quot;EVFold pipeline&quot; /&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Basically, given information about pairs of entities, 
there is a nice way to get rid of a lot of transitive interactions between pairs
to get a useful set of correlations between entities, so called couplings.
There are many places where you can apply this method,
but the couplings you get out of this system should represent something meaningful.
Otherwise, what you are doing is simply applying a method in an irrelevant manner.
For example, in protein world, the evolutionary couplings between residiues represent functional or structural constraints on those entities.
The tricky part is to find data sets where this method might provide you with meaningful results.&lt;/p&gt;

&lt;p&gt;I have known about this method for quite some time,
but didn’t have a data set to apply it to.
This is, of course, until recently when I had an epiphany about software systems.
While I was browsing the history of commits for a project of mine,
I realized that it holds a great deal of information about the software itself.&lt;/p&gt;

&lt;p&gt;For those who are not familiar with Version Contol Systems,
I strongly suggest you get yourself familiar with them.
In a nut-shell, these systems provides you the means to track changes you make to your source code
and to version them properly.
So when you look at the history of changes for a particular software project,
you see something like this:&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/img/evsrc-versioncontrol.png&quot; alt=&quot;Revision and changed files&quot; /&gt;&lt;/p&gt;

&lt;p&gt;This is, of course, an over-simplification, but the idea here is that you can take this information and create a matrix where you can mark (with 1) in which revision a particular file has changed.
And when you do this, you essentially get something similar to a sequence alignment:&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/img/evsrc-similarity.png&quot; alt=&quot;Matrix-like representation of revisions and alignment&quot; /&gt;&lt;/p&gt;

&lt;p&gt;Which is great, because you can use this piece information for inferring evolutionary coupled files in your software.
Once the data is in correct format, the inference can easily be done using one of &lt;em&gt;de facto&lt;/em&gt; R packages, e.g. &lt;a href=&quot;http://cran.r-project.org/web/packages/corpcor/index.html&quot;&gt;corpcor&lt;/a&gt;.
But what does it mean for source files to be coupled in a software evolution?&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/img/evsrc-coupling.png&quot; alt=&quot;Coupled files&quot; /&gt;&lt;/p&gt;

&lt;p&gt;Turns out that there is a huge body of literature on this,
but for those are curious, here is a way to learn more about this on Google Scholar: &lt;a href=&quot;http://scholar.google.com/scholar?hl=en&amp;amp;q=software+evolution+co-change&amp;amp;btnG=&amp;amp;as_sdt=1%2C33&amp;amp;as_sdtp=&quot;&gt;software evolution co-change&lt;/a&gt;.
In short, file couplings are often attributed to bad software design
and represent pieces of code that are either duplicated at some point in the history
or are not modular.
If so, then:&lt;/p&gt;

&lt;div class=&quot;quote&quot;&gt;
Evolutionary couplings between files can reveal poor design choices in software architecture.
&lt;/div&gt;

&lt;p&gt;To test this idea, I built this really simple pipeline as a proof of concept:&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/img/evsrc-pipeline.png&quot; alt=&quot;evSrc pipeline&quot; /&gt;&lt;/p&gt;

&lt;p&gt;This whole pipeline might seem like a big deal,
but it is actually a few lines of code that you can find here on GitHub: &lt;a href=&quot;https://github.com/armish/evsrc&quot;&gt;armish/evsrc&lt;/a&gt;.
It basically:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;clones a repository from GitHub into a local folder;&lt;/li&gt;
  &lt;li&gt;extracts information about changed files in each revision;&lt;/li&gt;
  &lt;li&gt;passes this information to an R-script that infers the couplings;&lt;/li&gt;
  &lt;li&gt;visualizes and outputs all the couplings with different specificity.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;People, of course, have been fancying with this idea for a long time,
but as far as I can tell nobody has approached this problem from this perspective.
To see if the inference step makes sense and whether it produces reasonable results,
I ran this pipeline first on &lt;a href=&quot;https://github.com/angular/angular.js&quot;&gt;Angular.js&lt;/a&gt;:&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/img/evsrc-angularjs.png&quot; alt=&quot;Angular.js file couplings&quot; /&gt;&lt;/p&gt;

&lt;p&gt;and then on &lt;a href=&quot;https://github.com/ipython/ipython&quot;&gt;iPython&lt;/a&gt;:&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;/img/evsrc-ipython.png&quot; alt=&quot;iPython file couplings&quot; /&gt;&lt;/p&gt;

&lt;p&gt;I did this because I knew that both pieces of software are well-maintained and well-designed;
so the pairs I get back, if any, should be easier to interpret.
In Angular.js, the approach revealed cute couplings, for example a coupling between &lt;a href=&quot;https://github.com/angular/angular.js/blob/master/images/css/arrow_left.gif&quot;&gt;left arrow image&lt;/a&gt; and &lt;a href=&quot;https://github.com/angular/angular.js/blob/master/images/css/arrow_right.gif&quot;&gt;right arrow image&lt;/a&gt;.
This makes sense, since every time one of them gets updated, the other one should also be taken care of.
In many frameworks, I saw people resolving this kind of issues simply by combining these two into a single sprite sheet.&lt;/p&gt;

&lt;p&gt;It also turns out that in both pieces of software, a majority of the couplings are attributable to &lt;a href=&quot;http://en.wikipedia.org/wiki/Test-driven_development&quot;&gt;Test Driven Design&lt;/a&gt;,
where a source file is coupled to its test.
So these are apparently false-positives I should take care of in the next version of the pipeline.
Beside these, I also get a handful of suspicous couplings that seem to be due bad design,
but they require another round of investigation before saying something concerete about them.&lt;/p&gt;

&lt;p&gt;In summary, infering evolutionary couplings seems to work fine in software development
and it seems to provide interesting information about the design of the software.
But to turn this evSrc project into a useful tool,
I first need to:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;find examples of badly-designed software projects&lt;/li&gt;
  &lt;li&gt;handle exceptional couplings due to good design (e.g. TDD)&lt;/li&gt;
  &lt;li&gt;improve the scripts so people can run it more easily&lt;/li&gt;
  &lt;li&gt;turn this into a web-based tool where people can simply submit jobs and get interactive networks out of it&lt;/li&gt;
  &lt;li&gt;think about ways to integrate this with &lt;a href=&quot;http://en.wikipedia.org/wiki/Integrated_development_environment&quot;&gt;Integrated Development Environments&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;…&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I might or might not be able to get to these items in the near future,
but at this point, any feedback is more than welcome.
Know papers, tools, approaches similar to this?
Or want to help me finalize this?
Then do let &lt;a href=&quot;http://ergoso.me/general/2014/01/06/about.html&quot;&gt;me&lt;/a&gt; know :)&lt;/p&gt;

</content>
 </entry>
 
 <entry>
   <title>The Road to Independence - Part III: Research proposal and CV</title>
   <link href="https://ergoso.me/career/2014/11/13/the-road-to-independence-research-proposal-and-cv.html"/>
   <updated>2014-11-13T20:00:01+00:00</updated>
   <id>https://ergoso.me/career/2014/11/13/the-road-to-independence-research-proposal-and-cv</id>
   <content type="html">&lt;p&gt;In &lt;a href=&quot;http://ergoso.me/career/2014/10/25/the-road-to-independence-what-it-takes.html&quot;&gt;my earlier post&lt;/a&gt;,
I talked about what it takes to apply to independent fellowship positions.
This is the third post in &lt;a href=&quot;http://ergoso.me/career/2014/10/16/the-road-to-independence-independent-postdoc-fellowships.html&quot;&gt;the series&lt;/a&gt;
and I am going to talk about my experience with coming up with a reasonable research proposal and CV.&lt;/p&gt;

&lt;h2 id=&quot;curriculum-vitae-cv&quot;&gt;Curriculum Vitae (C.V.)&lt;/h2&gt;
&lt;p&gt;When I started working on my application package,
the first thing I did was to update my CV.
I have previously &lt;a href=&quot;http://ergoso.me/lasker/essay/rejection/2014/08/25/discounted-tickets-for-science-education.html&quot;&gt;applied to a few scholarships&lt;/a&gt; earlier in my graduate life,
but each of them required a different format/style.
And to be honest, it was not my best work,
because most of these previous submissions were done at the very last minute, leaving room for a lot of errors and typos.
So I was not happy about the current state of the CV
and went looking for new, fresh examples to base my new one on.&lt;/p&gt;

&lt;p&gt;The one I liked the most was &lt;a href=&quot;http://www.choderalab.org/&quot;&gt;John Chodera&lt;/a&gt;’s amazingly clean and explanatory CV.
I knew this particular one, 
because we happened to be in the same department
and he posted a link to &lt;a href=&quot;https://github.com/jchodera/latex-cv&quot;&gt;the repository of his own CV&lt;/a&gt; in our internal e-mail list.
I then went on &lt;a href=&quot;http://drbecca.scientopia.org/tt-job-search-advice-aggregator/&quot;&gt;the advice aggregator&lt;/a&gt; and read every single post on how to prepare CVs.&lt;/p&gt;

&lt;p&gt;After all this hassle, I had a good idea about how to prepare my own CV
and the things I would like to highlight in it.
For the curious, here is &lt;a href=&quot;http://arman.aksoy.org/AksoyBA_CV.pdf&quot;&gt;the latest version&lt;/a&gt; of mine.&lt;/p&gt;

&lt;p&gt;But in short, here are the advices I took into consideration when coming up with mine:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Put links to your papers so people can quickly access them if they are reading your application on their computer&lt;/li&gt;
  &lt;li&gt;Add your e-mail address and the URL for your personal web site where people will be able to find up-to-date information about your projects and your CV&lt;/li&gt;
  &lt;li&gt;Group your publications into categories so that the related ones are closer to each other&lt;/li&gt;
  &lt;li&gt;At the top of your publication list, put some summary statistics (no, not your h-index) so that people will know what is waiting for them&lt;/li&gt;
  &lt;li&gt;Get rid of anything that is irrelevant to your applications (your GPAs, GRE scores, your not-so-useful undergraduate internships)&lt;/li&gt;
  &lt;li&gt;Do not list every single poster you presented unless it is necessary to do so or you received an award for one&lt;/li&gt;
  &lt;li&gt;Instead of over-promising and under-delivering, under-promise but over-deliver&lt;/li&gt;
  &lt;li&gt;Keep non-academic audience in mind and do not over-optimize for academic purposes&lt;/li&gt;
  &lt;li&gt;Use a clean, readable font face and do not over pack your text&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h2 id=&quot;research-proposal&quot;&gt;Research Proposal&lt;/h2&gt;
&lt;p&gt;This was one really hard to prepare.
First of all, you don’t find too many examples online
and even if you do, they are usually not good for your particular field.
Second, other than the cap on the number of pages, there is no clear formating guidelines for this document.
Third, it takes a lot of iterations to come up with a descent one.&lt;/p&gt;

&lt;p&gt;But don’t you worry!
Turns out that everybody has their own way of writing this piece of material and
there is no real consensus on what this document should contain and how it should be written.
But do your best in terms of making it easy for a committee member to understand what you are trying to tell.
The best way to do that is to imagine yourself sitting in a room with hundreds of applications 
and taking a look at a particular proposal for a few minutes to decide whether it is good or bad.&lt;/p&gt;

&lt;p&gt;So do take your time to go over &lt;a href=&quot;http://jabberwocky.weecology.org/2012/08/10/a-list-of-publicly-available-grant-proposals-in-the-biological-sciences/&quot;&gt;this amazing list of grant proposals&lt;/a&gt; (aggregated by &lt;a href=&quot;http://whitelab.weecology.org/&quot;&gt;Ethan White&lt;/a&gt;) and &lt;a href=&quot;http://drbecca.scientopia.org/tt-job-search-advice-aggregator/&quot;&gt;the advices&lt;/a&gt;.
If any, take a look at either &lt;a href=&quot;http://www.ethanperlstein.com/my-postdoc-research-proposal/&quot;&gt;Ethan Perlstein’s&lt;/a&gt; and &lt;a href=&quot;http://ged.msu.edu/downloads/2013-research.pdf&quot;&gt;Titus Brown’s&lt;/a&gt; statements, both of which were accepted, so sucessful examples.&lt;/p&gt;

&lt;p&gt;Besides these publicly available ones, I found out that most fellows are OK with the idea of sharing their proposal if you ask them.
Go ahead and see if any of the fellows that are in a program of interest to you are in your network,
and ask them if they are willing to share their proposal with you.
Of course, I still don’t know if mine is a good one or not, 
but inspired by all those great folks,
I made &lt;a href=&quot;https://github.com/armish/research-proposal&quot;&gt;my proposal public on GitHub&lt;/a&gt; (forks are welcome!).&lt;/p&gt;

&lt;iframe src=&quot;http://wl.figshare.com/articles/1239172/embed?show_title=1&quot; width=&quot;100%&quot; height=&quot;600&quot; frameborder=&quot;0&quot;&gt;&lt;/iframe&gt;

&lt;p&gt;Read these as if you are one of the members of the search committee.
Do you like it? Do you understand it at a basic level? Is it exciting?
If so, how come? What made you like that particular proposal?
Consider all these and adopt the things you liked and avoid the things you didn’t like when writing yours.&lt;/p&gt;

&lt;p&gt;First, the more you read these examples, the more comfortable you will be when writing up.
So try to read as many examples as possible.
Second, my ultimate advice about writing your research proposal is to start writing as early as you can
and do ask for feedback from people around you.
I e-mailed quite a few to ask for feedback and only half of them had the time to comment on the proposal, so plan accordingly.
And those comments I got so far, were incredibly useful and helped me improve my proposal a lot.&lt;/p&gt;

&lt;p&gt;But keep in mind that you should not take every single one of these comments seriously and change your proposal to make everyone happy.
&lt;a href=&quot;http://www.thespectroscope.com/read/seek-advice-and-learn-to-ignore-it-by-lenny-teytelman-240&quot;&gt;Seek advice and learn to ignore it&lt;/a&gt;.
The way I handled all the feedback I got was to categorize things people pointed out in my proposal
and fix the things that bothered more than three people:&lt;/p&gt;

&lt;blockquote class=&quot;twitter-tweet&quot; lang=&quot;en&quot;&gt;&lt;p&gt;so I heard from 5 persons regarding my research proposal and here is a summary of the feedback I got so far: &lt;a href=&quot;http://t.co/Gngkm6YTKT&quot;&gt;pic.twitter.com/Gngkm6YTKT&lt;/a&gt;&lt;/p&gt;&amp;mdash; B. Arman Aksoy (@armish) &lt;a href=&quot;https://twitter.com/armish/status/528984135340404738&quot;&gt;November 2, 2014&lt;/a&gt;&lt;/blockquote&gt;
&lt;script async=&quot;&quot; src=&quot;//platform.twitter.com/widgets.js&quot; charset=&quot;utf-8&quot;&gt;&lt;/script&gt;

&lt;p&gt;Here are some words of wisdom from various people I contacted about writing a research proposal:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Clearly state your past work, future work and the motivation behind it&lt;/li&gt;
  &lt;li&gt;Guide readers with titles/sections and always provide quick summaries for skim readers&lt;/li&gt;
  &lt;li&gt;Do not use small fonts and do not exceed the page limits&lt;/li&gt;
  &lt;li&gt;Be concise and clear about what you want to say (have no mercy for wordy paragraphs)&lt;/li&gt;
  &lt;li&gt;Do not say negative things in your proposal and do not bull shit others’ work&lt;/li&gt;
  &lt;li&gt;Provide some reasoning about why you are the correct person to do the things you are proposing&lt;/li&gt;
  &lt;li&gt;Do not be too vague about your ideas (I will cure cancer) and do not provide too much details (I’m going to use this kit [catalog number])&lt;/li&gt;
  &lt;li&gt;Do not over use bold/italic fonts&lt;/li&gt;
  &lt;li&gt;If you have enough space, throw a conceptual figure in there&lt;/li&gt;
  &lt;li&gt;Do not say things like “after I graduate …” (your career is continious and your either pre- or post-doctoral work is still your work)&lt;/li&gt;
  &lt;li&gt;Do not treat this as a post-doc proposal where you are simply going to extend your earlier work. Fellows are encouraged and expected to establish and lead a new research program within the departments (so I was told)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you are like me and English is your secondary language, then I strongly suggest you to ask help from a science writer on editing your final draft.
This will not only help you get rid of too scientific terminology, but also will improve the clarity of your message.
If you don’t know any science writers, then you might think about getting help from one of the professional editing services out there.&lt;/p&gt;

&lt;h3 id=&quot;wrapping-up&quot;&gt;Wrapping up&lt;/h3&gt;
&lt;p&gt;These two are really important parts of the application package, but they are not the final determinants.
You still need to tweak and customize some of these to be able to start submitting your application and maximize their impact on the search committee.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Coming up next&lt;/strong&gt;: Putting your application package together.&lt;/p&gt;
</content>
 </entry>
 
 <entry>
   <title>The Road to Independence - Part II: What it takes to apply for fellowships</title>
   <link href="https://ergoso.me/career/2014/10/25/the-road-to-independence-what-it-takes.html"/>
   <updated>2014-10-25T17:00:01+00:00</updated>
   <id>https://ergoso.me/career/2014/10/25/the-road-to-independence-what-it-takes</id>
   <content type="html">&lt;p&gt;As I wrote earlier, 
&lt;a href=&quot;http://ergoso.me/career/2014/10/16/the-road-to-independence-independent-postdoc-fellowships.html&quot;&gt;I am interested in becoming an early independent fellow&lt;/a&gt;
and started a blog series about it.
This is the second post in the series.&lt;/p&gt;

&lt;p&gt;In this post, I am going to talk about what it takes to apply for such a fellowship.
And for those of you who want me to cut to the chase, 
here are things you will be needing for application:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Lots of time (and I mean lots) for finding such fellowship opportunities online&lt;/li&gt;
  &lt;li&gt;Lots of time for e-mailing program coordinators to get the details right&lt;/li&gt;
  &lt;li&gt;A short research proposal that justifies your request for independence&lt;/li&gt;
  &lt;li&gt;At least three strong recommendation letters&lt;/li&gt;
  &lt;li&gt;A good publication record for the prestigious ones&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Let me talk about these one by one and share my experience with you for each of them separately.&lt;/p&gt;

&lt;h2 id=&quot;finding-postdoc-fellowship-opportunities&quot;&gt;Finding postdoc fellowship opportunities&lt;/h2&gt;
&lt;p&gt;&lt;a href=&quot;http://ergoso.me/career/2014/10/16/the-road-to-independence-independent-postdoc-fellowships.html&quot;&gt;As I mentioned&lt;/a&gt; earlier, I got to know about the existence of such fellowships just by chance.
I happened to be talking to people and stalking their CVs online,
which introduced me to one of these fellowships just by coincidence.&lt;/p&gt;

&lt;p&gt;I usually take a quick look at the job postings whenever I have a chance
and try to get a sense of what is out there waiting for me;
but not once, in my humble 5-year graduate school years, I saw a posting for such a fellowship.
And this is a fact you should be aware of: these fellowships are not advertised widely for some reason.
So you have to be the one who is chasing after such opportunities.&lt;/p&gt;

&lt;p&gt;And as in every chasing game, it takes a lot of effort and time to get a complete listing of these.
There was a relatively old article in Cell about these fellowships, titled “&lt;a href=&quot;http://www.ncbi.nlm.nih.gov/pubmed/17382872&quot;&gt;Superpostdocs reach for the stars&lt;/a&gt;”,
which was a really good start for me, but it turned out that many fellowships mentioned in that article are no more accepting new applications.
Googling relevant keywords, talking to people and e-mailing departments of interest to me helped a lot to nail some of them down,
and for the curious, here is the list of applications I will be applying to:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Princeton Lewis-Sigler Fellowship (&lt;a href=&quot;http://www.princeton.edu/genomics/lewis-sigler-fellows/&quot;&gt;info&lt;/a&gt;)&lt;/li&gt;
  &lt;li&gt;CSHL Fellows Program (&lt;a href=&quot;http://www.cshl.edu/Research/CSHL-Fellows-Program.html&quot;&gt;info&lt;/a&gt;)&lt;/li&gt;
  &lt;li&gt;MIT Whitehead Fellow Program (&lt;a href=&quot;http://wi.mit.edu/people/fellows&quot;&gt;info&lt;/a&gt;)&lt;/li&gt;
  &lt;li&gt;Rockefeller BioPhysics Fellow Program (&lt;a href=&quot;http://uqbar.rockefeller.edu/fellows.html&quot;&gt;info&lt;/a&gt;)&lt;/li&gt;
  &lt;li&gt;Max-Planck Research Group Leader Program (&lt;a href=&quot;http://www.mpg.de/mprg_apply&quot;&gt;info&lt;/a&gt;)&lt;/li&gt;
  &lt;li&gt;NIH Early Independence Award (&lt;a href=&quot;http://commonfund.nih.gov/earlyindependence/index&quot;&gt;info&lt;/a&gt;)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I honestly think that we need a crowd-sourcing approach to create a complete listing of such opportunities to get the word out as much as possible for the soon-to-be-graduates.
But that requires another blog post on its own.&lt;/p&gt;

&lt;h2 id=&quot;getting-the-details-right&quot;&gt;Getting the details right&lt;/h2&gt;
&lt;p&gt;One thing that bothered me a lot about these fellowships is that every single one of them seem to have a different application system to them. 
Some of them require your mentor to first nominate you to the program before you can apply,
some of them want you to submit everything as a single application package via e-mail,
and some follow a typical grant/faculty application procedure where they have a proper electronic system to manage things.
Making sure to meet all the deadlines, not miss anything important and not filing something that is completely out-of-context is hard,
because most of these information pieces are not available online.
So get prepared for sending many e-mails to program coordinators
and in some cases also for chasing the right contact person to send the e-mail to.&lt;/p&gt;

&lt;p&gt;In my case, it took on average two e-mails back and forth to get all details I needed for these applications.
To keep track of all these conversations, I suggest you get yourself familiar with task manager services (my favorite one is &lt;a href=&quot;https://www.rememberthemilk.com/&quot;&gt;RememberTheMilk&lt;/a&gt;).
I also find myself visiting web-pages for these fellowships quite frequently to see when the new application period will be open
and unfortunately, not all of them have RSS feeds to them.
To keep track of these web page changes, I ended up setting up alerts on &lt;a href=&quot;https://www.changedetection.com/&quot;&gt;ChangeDetection&lt;/a&gt; web service
and it works beautifully and takes a huge burden off of your shoulders.&lt;/p&gt;

&lt;h2 id=&quot;at-least-three-recommendation-letters&quot;&gt;At least three recommendation letters&lt;/h2&gt;
&lt;p&gt;This one is really challenging.
You see, most graduates students work in my field work in isolation and work with their advisor on a project.
Under normal circumstances, this means that by default a graduate student has two secured reference letters: 
one from her graduate school advisor and one from her previous advisor.
The latter is somewhat questionable, because Ph.D. programs usually takes 5-6 years on average
and during all these years, many graduate students diverge from the field that they received their degree in.
And this means that the former advisor’s letter might not count as strong or informative as your new advisor’s letter.&lt;/p&gt;

&lt;p&gt;Thanks to my &lt;a href=&quot;http://www.triiprograms.org/cbm/&quot;&gt;tri-institutional graduate program&lt;/a&gt;, I got to spend a year in a different campus and had the chance to work in another lab for a whole year (+1).
I also happened to work for a multi-institutional project and interacted with another PI quite frequently (+1).
And I have no serious issues with my current advisor (+1).
In this sense, I was lucky to secure these additional support letters but not everybody is fortunate enough to do so.&lt;/p&gt;

&lt;p&gt;So my advice to you, dear student friend who happened to be reading this post and looking for humble advice, is as follows:
participate in collaborative projects as much as possible and get to know people.
If you are considering becoming independent early on, then start acting independently as early as possible.
If your lab collaborates a lot, then be part of at least one collaborative projects;
if not, then think about starting one on your own.&lt;/p&gt;

&lt;hr /&gt;

&lt;p&gt;Let me pause here and leave the next two important things for another blog post before this one gets too long and boring.
But the bottom line of this post is that these applications do take time 
and can be considered as a part-time job.
For those of you who are already overwhelmed with other things, I suggest you think twice before going down this path.
If you are seriously thinking about applying for these fellowships,
then plan ahead and try to dedicate a whole month to get these applications out of your way.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Coming up next&lt;/strong&gt;: &lt;a href=&quot;http://ergoso.me/career/2014/11/13/the-road-to-independence-research-proposal-and-cv.html&quot;&gt;Preparing a CV and a research proposal&lt;/a&gt;.&lt;/p&gt;
</content>
 </entry>
 
 <entry>
   <title>The Road to Independence - Part I: Applying for Independent Postdoc Fellowships</title>
   <link href="https://ergoso.me/career/2014/10/16/the-road-to-independence-independent-postdoc-fellowships.html"/>
   <updated>2014-10-16T12:00:01+00:00</updated>
   <id>https://ergoso.me/career/2014/10/16/the-road-to-independence-independent-postdoc-fellowships</id>
   <content type="html">&lt;p&gt;&lt;em&gt;I am a fifth-year graduate student … and considering applying for such and such fellowships…&lt;/em&gt; 
So starts my e-mail messages to program coordinators nowadays.
I am now officially a senior student in the Ph.D. program and on average, this is the year I should wrap things up and start thinking about &lt;em&gt;the next step&lt;/em&gt;.
And I am.
And it is taking a lot of time and effort to do these things.
And there is definitely no good guide for what I am trying to do.
This is exactly why I decided to go transparent about this process
and blog about certain aspects of applying for an independent postdoc fellowship, right off of graduate school.&lt;/p&gt;

&lt;p&gt;Everything started with me coming across &lt;a href=&quot;http://www.ethanperlstein.com/&quot;&gt;Ethan Perlstein&lt;/a&gt;’s &lt;a href=&quot;http://www.rockethub.com/projects/11106-crowdsourcing-discovery&quot;&gt;Crowdsourcing Discovery&lt;/a&gt; project.
I really liked the idea of supporting a science project like this 
and getting a 3D-printed methamphetamine molecule as a bonus.
So I backed the project.
Fast-forward a few months and I received an e-mail from Ethan about his being in the city 
and checking the possibility for a meet up so that he could give me the gift.
This was extremely timely, 
because we happened to be looking for new presenters for a science discussion club we had been running at our lab.
So I replied back to him asking whether he was willing to present his work
and he agreed to do so.
And right before I sent &lt;a href=&quot;http://www.ethanperlstein.com/wp-content/uploads/2012/06/Perlstein_CV.pdf&quot;&gt;his CV&lt;/a&gt; to our group for the introduction,
I found myself googling &lt;a href=&quot;http://www.princeton.edu/genomics/lewis-sigler-fellows/&quot;&gt;Lewis-Sigler Fellowship&lt;/a&gt; to see what kind of a fellowship it was.&lt;/p&gt;

&lt;p&gt;At that point, my third year in the graduate program, 
I absolutely was not even thinking about graduation, leave alone where to apply next.
I did know that I wanted to do a post-doc and stay in academia for good,
but before then, I didn’t know about the possibility of becoming an independent post-doc.
So I started asking people about the advantages and disadvantages of these fellowships.
After a fruitful discussion with &lt;a href=&quot;https://twitter.com/lteytelman&quot;&gt;Lenny Teytelman&lt;/a&gt;,
I made up my mind and &lt;a href=&quot;https://www.pubchase.com/career/question/from-a-naive-point-of-view-independent-post-doc-fellowships-196&quot;&gt;posted a question&lt;/a&gt; about these fellowships on &lt;a href=&quot;https://www.pubchase.com/career/question/from-a-naive-point-of-view-independent-post-doc-fellowships-196&quot;&gt;PubChase&lt;/a&gt;
and got back really informative answers from former/current fellows.&lt;/p&gt;

&lt;p&gt;After carefully considering these options and discussing future plans with &lt;a href=&quot;https://twitter.com/PINPINPINN&quot;&gt;my dearest wife&lt;/a&gt;,
we decided that these fellowships were definitely worth applying for
before going for a traditional postdoc position.
Giving this decision was already a major thing,
but preparing the application turned out to be much harder than I anticipated.&lt;/p&gt;

&lt;p&gt;As a graduate student who &lt;a href=&quot;http://ergoso.me/lasker/essay/rejection/2014/08/25/discounted-tickets-for-science-education.html&quot;&gt;previously applied for smallish fellowships but got rejected from all&lt;/a&gt;,
and as a student who doesn’t have much of an experience in terms of putting in successful grant applications,
you can probably feel my pain about working on a killer fellowship application.
And add my being international to the equation, which reduces the list of possible positions that I can apply to down to a handful.&lt;/p&gt;

&lt;p&gt;Sure, there are really great resources out there talking about how funding mechanism works,
how to write successful grants
and how to better apply for faculty positions (see &lt;a href=&quot;http://drbecca.scientopia.org/tt-job-search-advice-aggregator/&quot;&gt;Dr. Becca’s advice aggregator&lt;/a&gt; for more).
There are even more of these on &lt;a href=&quot;http://startupclass.samaltman.com/&quot;&gt;the start-up culture&lt;/a&gt;,
but it turns out that there is no such resource for people like me.&lt;/p&gt;

&lt;p&gt;I am also a huge fan of &lt;a href=&quot;https://twitter.com/abexlumberg&quot;&gt;Alex Blumberg&lt;/a&gt;’s &lt;a href=&quot;http://hearstartup.com/&quot;&gt;incredible podcast&lt;/a&gt; on starting a business and
as I listen to him talking about his experience, I realized many similarities between his and mine.
This got me thinking about becoming transparent about this whole application process
and document things along the way.
I might not get any of the fellowships,
so I don’t know if this blog series is going to be about a successful or a failed application;
but, I am hoping that it will prove useful for somebody who is thinking about becoming an independent postdoc.&lt;/p&gt;

&lt;blockquote class=&quot;twitter-tweet&quot; lang=&quot;en&quot;&gt;&lt;p&gt;&lt;a href=&quot;https://twitter.com/armish&quot;&gt;@armish&lt;/a&gt; &lt;a href=&quot;https://twitter.com/eperlste&quot;&gt;@eperlste&lt;/a&gt; do it! ; )&lt;/p&gt;&amp;mdash; StartUp (@podcaststartup) &lt;a href=&quot;https://twitter.com/podcaststartup/status/520381437410050048&quot;&gt;October 10, 2014&lt;/a&gt;&lt;/blockquote&gt;
&lt;script async=&quot;&quot; src=&quot;//platform.twitter.com/widgets.js&quot; charset=&quot;utf-8&quot;&gt;&lt;/script&gt;

&lt;p&gt;&lt;strong&gt;Coming up next&lt;/strong&gt;: &lt;a href=&quot;http://ergoso.me/career/2014/10/25/the-road-to-independence-what-it-takes.html&quot;&gt;thoughts about the application packages, recommendation letters and networking&lt;/a&gt;.&lt;/p&gt;
</content>
 </entry>
 
 <entry>
   <title>Of Synthetic Lethals and Vulnerabilities in Cancer</title>
   <link href="https://ergoso.me/cancer/therapy/vulnerability/personalized/precision/medicine/2014/09/02/of-vulnerabilities-and-synthetic-lethalities.html"/>
   <updated>2014-09-02T16:00:01+00:00</updated>
   <id>https://ergoso.me/cancer/therapy/vulnerability/personalized/precision/medicine/2014/09/02/of-vulnerabilities-and-synthetic-lethalities</id>
   <content type="html">&lt;h2 id=&quot;the-dream&quot;&gt;The Dream&lt;/h2&gt;
&lt;p&gt;Imagine a cancer patient walking into the clinic to learn her therapy options.
Her tumor sample was collected some while ago and now it is time to hear the lab results.
Results, of course, are from a comprehensive -omics analysis of the tumor sample
and are tailored for her.
As soon as they pop-up on the physician’s screen,
the physician takes a careful look at them and prescribes a drug for the patient suggested by the analysis.
The drug is no common cancer drug; 
it is a drug that has been on the market for some while but is being used for treating another disease.
It turns out, from the -omics analysis of this patient, this drug should magically kill only the tumor cells and hence should not have any side effects for the patient.
Patient starts on the therapy, and as expected, the tumor starts shrinking and eventually disappears.
The patient goes on with her life normally during/after the therapy.
She is happy to hear the good news and leaves the clinic with a smile.
As she leaves, another cancer patient walks into the clinic to learn his personalized therapy options…&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;http://bioinformatics.oxfordjournals.org/content/30/14/2051/F5.expansion.html&quot;&gt;&lt;img src=&quot;/img/personalized-cancer-therapy.png&quot; alt=&quot;Personalized cancer therapy&quot; /&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The above scenario, although it may sound as too good to be true, is not completely impossible with the latest technology and our current understanding of the cancer genome.
In fact, many universities and instutions are already moving in this direction
and started thinking seriously about the way they approach this personalized medicine idea.
And in the meantime, the need for a computational tool or a set of tools that will give us some idea about which patient should be treated with which drug is becoming more apparent every single day.&lt;/p&gt;

&lt;p&gt;Treating a tumor that has HER2-amplified with trastuzumab is a no-brainer.
Nor it is for a BRAF V600E mutant patient with vemurafenib.
But what if the patient doesn’t have any of those &lt;em&gt;actionable&lt;/em&gt; alterations?
How do you decide on the magical drug to use in the treatment?
Computational approaches are to rescue!&lt;/p&gt;

&lt;h2 id=&quot;the-concept&quot;&gt;The Concept&lt;/h2&gt;
&lt;p&gt;But before going into these computational approaches,
let’s take a moment to step back and take a look at a few other concepts.
&lt;em&gt;Synthetic lethality&lt;/em&gt;, for example, is an important one.
If gene &lt;em&gt;A&lt;/em&gt; and &lt;em&gt;B&lt;/em&gt; are synthetic lethal pairs,
it simply means that cells can do fine without gene A or without gene B but not without both.
Then there are &lt;em&gt;essential&lt;/em&gt; genes.
Essentiality is highly context-specific, 
meaning that a gene might be essential for cells when a specific condition is met.
A trivial example to this are synthetic lethal pairs,
where one of these genes become essential when the other one is knocked out.
That is, gene A is essential when B is already knocked-out.
Sometimes you can get more complex patterns,
where, for example, gene C might become essential when gene A is over-active.
It is also worth noting that these two terms can be, from time to time, be used interchangeably.&lt;/p&gt;

&lt;p&gt;These concepts are even more interesting when we are talking about cancer,
because cancer cells are a mess with their instable genomes.
In the process of a normal cell to become a cancer cell,
a cell accumulates lots of random alterations.
And not surprisingly, these alterations sometimes disable one of the partners of a synthetic lethal pair.
Or due to a gain-of-function alteration, 
these might cause some other genes to become essential for the cancer cell.
Cells can tolerate these events considerably well until either
the cell loses the partner in the synthetic lethal pair or
we perturb this cell and hit the other partner, for example, by a targeted drug.
The former, we don’t see much in the data sets,
because those cells without its essential genes are selected against and immediately disappear from the tumor.
The latter, it is interesting, because these are the events that create &lt;em&gt;vulnerabilities&lt;/em&gt; in cancer cells.
Since such vulnerabilities are restricted to the cancer cells,
they also create an opportunity for us to exploit them as a therapeutic approach.
You can think of this as cancer cells losing their gene A 
and then us hitting gene B with a drug, 
hence selectively killing these cancer cells.
And be aware that, 
since gene A is lost only in cancer cells, 
normal cells in the body will still do fine in this case, i.e. the ideal therapy with minimum toxicity to the host.&lt;/p&gt;

&lt;h2 id=&quot;the-approach&quot;&gt;The Approach&lt;/h2&gt;
&lt;p&gt;Now that we know what the problem (cancer) and even a possible solution to it (therapeutic vulnerability),
we can start exploring the approaches to reveal clinically relevant synthetic lethal gene pairs.&lt;/p&gt;

&lt;p&gt;Let’s start with a trivial example
and assume the following: 
&lt;em&gt;every gene pair represents a synthetic lethality (SL) group&lt;/em&gt;.
Assuming that we know about &lt;a href=&quot;http://www.genenames.org/&quot;&gt;40000 gene symbols&lt;/a&gt;, 
this extreme approach leads to around 160 &lt;em&gt;million&lt;/em&gt; potential SL pairs.
This is just a way of suggesting SL pairs, but is it helpful? 
Not really so.&lt;/p&gt;

&lt;p&gt;Instead, let’s be more realistic and work with &lt;em&gt;in vitro&lt;/em&gt; models: cell lines.
We can, for example, take a bunch of cell lines for which we know the genetic alterations,
and try to hit every single gene they have one by one and observe the effects.
The genes, when hit, that kill some of the cell lines are interesting,
because they are potentially essential to those cells
and represent therapeutic targets.
This is not a bad idea at all and
in fact, there are really nice studies (&lt;a href=&quot;http://www.broadinstitute.org/achilles&quot;&gt;Project Achilles&lt;/a&gt; and &lt;a href=&quot;http://cancerdiscovery.aacrjournals.org/content/2/2/172.short&quot;&gt;Marcotte &lt;em&gt;et al.&lt;/em&gt;&lt;/a&gt;) out there that conduct this type of a screen on cell lines.
The technology of choice for these experiments is usually pooled-shRNAs combined with deep-sequencing.
You can also swap shRNAs with drugs and get a similar data set centered around compounds (e.g. &lt;a href=&quot;http://www.broadinstitute.org/ccle/&quot;&gt;CCLE&lt;/a&gt;, &lt;a href=&quot;http://www.cancerrxgene.org/&quot;&gt;CancerRxGene&lt;/a&gt; or &lt;a href=&quot;http://www.broadinstitute.org/ctrp/&quot;&gt;CTRP&lt;/a&gt;).
The problem with this type of an experimental approach, however, is that it is really hard to scale things up,
because these experiments are expensive and labor-intensive.
And cell line pools do not always represent the full patient diversity, hence are under-powered.&lt;/p&gt;

&lt;p&gt;Yet another way to approach this problem is to model cells, that is working computationally on a systems-level.
You can, for example, learn how normal and cancer cells operate on a metabolism level
and observe what happens to these cells if you hit a gene in both.
The ideal hit disrupts the metabolism in cancer cells bad enough to kill them or stop them proliferate,
but does not effect the normal cells.
This is a feasible and reasonable approach (see &lt;a href=&quot;http://msb.embopress.org/content/7/1/501&quot;&gt;Folger &lt;em&gt;et al.&lt;/em&gt;&lt;/a&gt;),
but for this, you need to rely on generalized normal/cancer models that cover a majority of the patients but not all of them.&lt;/p&gt;

&lt;p&gt;Or you can leverage what we know about metabolism so far in a qualitative manner
and assume that cells need all these reactions to function correctly.
Then you can take a look at all enzymes that catalyze the same reaction in a cell, so-called &lt;em&gt;isoenzymes&lt;/em&gt;,
and see if a cancer cell lacks one particular isoenzyme, where the other partner can be targeted with a drug.
The idea here is that losing a metabolic enzyme creates a vulnerability for the cancer cells,
and you can exploit this vulnerability by perturbing these cells with a drug to inhibit an isoenzyme partner.
We recently implemented this method and &lt;a href=&quot;http://bioinformatics.oxfordjournals.org/content/30/14/2051&quot;&gt;report our findings here&lt;/a&gt;.
The advantage of this approach is that it allows to nominate vulnerabilities and drugs associated with them in a &lt;a href=&quot;http://cbio.mskcc.org/cancergenomics/statius/&quot;&gt;personalized manner&lt;/a&gt;.
The disadvantage is that you rely on prior knowledge 
and not all isoenzymes that you infer from knowledge bases might represent SL gene groups.&lt;/p&gt;

&lt;p&gt;Finally, you can simply ignore the modeling and the prior knowledge parts,
and let the data speak for itself.
Given the amount of public data that we have,
an integrated analysis will tell us which genes are potentially SL pairs
–by looking at their co-alteration frequency and expression patterns–
and whether there is experimental data, e.g. shRNA screen results, to back this observation up.
As this really nice study by &lt;a href=&quot;http://dx.doi.org/10.1016/j.cell.2014.07.027&quot;&gt;Jerby-Arnon &lt;em&gt;et al.&lt;/em&gt;&lt;/a&gt; shows,
using these data is no easy task 
and requires combining many smart methods for proper integration and nomination.
But it seems doable and at the end provides us with a general network of all SL pairs.&lt;/p&gt;

&lt;h2 id=&quot;the-conclusion&quot;&gt;The Conclusion&lt;/h2&gt;
&lt;p&gt;No matter what kind of method or approach we adopt,
accumulating results from all these synthetic lethals and essentials are no doubt exciting.
With the increasing amount of data and new smart methods,
we will soon have our huge catalog of therapy options with specific contexts that they are expected to work.
And once we reach to that completeness,
we will be much closer to the level where we treat cancer as if it is a chronic disease.
And I am looking forward to those day
while trying to contribute to this ultimate aim as much as I can.&lt;/p&gt;

&lt;h2 id=&quot;the-rant&quot;&gt;The Rant&lt;/h2&gt;
&lt;p&gt;For those that are curious, here is me talking about &lt;a href=&quot;http://bioinformatics.oxfordjournals.org/content/30/14/2051&quot;&gt;our approach and results&lt;/a&gt; during the 3rd TCGA Scientific Symposium.&lt;/p&gt;

&lt;center&gt;
	&lt;iframe width=&quot;560&quot; height=&quot;315&quot; src=&quot;//www.youtube.com/embed/UvH3qRepw7Q&quot; frameborder=&quot;0&quot; allowfullscreen=&quot;&quot;&gt;&lt;/iframe&gt;
&lt;/center&gt;
</content>
 </entry>
 
 <entry>
   <title>Discounted Tickets for Science Education</title>
   <link href="https://ergoso.me/lasker/essay/rejection/2014/08/25/discounted-tickets-for-science-education.html"/>
   <updated>2014-08-25T13:00:01+00:00</updated>
   <id>https://ergoso.me/lasker/essay/rejection/2014/08/25/discounted-tickets-for-science-education</id>
   <content type="html">&lt;p&gt;I do terrible when it comes to application for fellowships.
I applied to seven different Ph.D. programs, but got officially accepted only by &lt;a href=&quot;http://www.triiprograms.org/cbm/&quot;&gt;one&lt;/a&gt; for which I had to appeal to their initial rejection.
I tried applying to &lt;a href=&quot;https://www.hhmi.org/programs/international-student-research-fellowships&quot;&gt;HHMI’s International Student Research Fellowship&lt;/a&gt;,
but &lt;a href=&quot;http://weill.cornell.edu/&quot;&gt;my institution&lt;/a&gt; did not nominate me for application.
I then applied to an internal fellowship at &lt;a href=&quot;http://www.mskcc.org/&quot;&gt;my other institution&lt;/a&gt; thinking that my chances would be higher given relatively smaller applicant pool.
I got rejected, but was &lt;em&gt;strongly encouraged&lt;/em&gt; to apply for next year’s fellowship.
After &lt;a href=&quot;https://www.pubchase.com/career/question/do-non-prestigious-pre-doctoral-fellowships-help-with-findin-162&quot;&gt;lots of thinking&lt;/a&gt;, I applied next year with high hopes and got rejected again.
I also took my chances at &lt;a href=&quot;https://www.facebook.com/weillcornellgradschool/posts/1480330052202268&quot;&gt;a research paper prize&lt;/a&gt; with &lt;a href=&quot;http://bioinformatics.oxfordjournals.org/content/30/14/2051&quot;&gt;my recently published work&lt;/a&gt;, but alas.&lt;/p&gt;

&lt;p&gt;So when I first heard about this year’s &lt;a href=&quot;http://www.laskerfoundation.org/programs/contest.htm&quot;&gt;Lasker Essay Contest&lt;/a&gt;,
I tried not to get so excited, but I made my mind and gave it a try. 
Today I got the news that the essay I submitted was not selected
and &lt;a href=&quot;https://twitter.com/armish/status/496657681772322816&quot;&gt;as I promised&lt;/a&gt;, I decided to make a blog post out of it.
The silver lining is that I now have yet another public file on &lt;a href=&quot;http://figshare.com/articles/Discounted_Tickets_for_Science_Education/1150282&quot;&gt;FigShare&lt;/a&gt;
and the visibility of this essay will hopefully be much higher with this post.&lt;/p&gt;

&lt;p&gt;For those of you that are curious, (&lt;em&gt;drum rolls please&lt;/em&gt;) &lt;a href=&quot;http://figshare.com/articles/Discounted_Tickets_for_Science_Education/1150282&quot;&gt;here&lt;/a&gt; is my essay on &lt;em&gt;innovative ways to build support and ensure funding for medical research&lt;/em&gt; – ah, the irony…:&lt;/p&gt;

&lt;iframe src=&quot;http://wl.figshare.com/articles/1150282/embed?show_title=0&quot; width=&quot;100%&quot; height=&quot;770&quot; frameborder=&quot;0&quot;&gt;&lt;/iframe&gt;

</content>
 </entry>
 
 <entry>
   <title>bioRxiv - The Case of Removed Papers</title>
   <link href="https://ergoso.me/biorxiv/preprint/paper/2014/07/07/biorxiv-unpublished-papers.html"/>
   <updated>2014-07-07T15:40:01+00:00</updated>
   <id>https://ergoso.me/biorxiv/preprint/paper/2014/07/07/biorxiv-unpublished-papers</id>
   <content type="html">&lt;p&gt;I love preprint services. 
I love the idea of archiving preprints, short-cutting the painful peer-review process and sharing the results in an easy and quick manner.
This is why for &lt;a href=&quot;http://www.biorxiv.org/content/early/2014/05/29/005686&quot;&gt;our latest paper&lt;/a&gt;, 
I insisted on experimenting with the idea and actually ended up posting the paper on &lt;a href=&quot;http://www.biorxiv.org/content/early/2014/05/29/005686&quot;&gt;bioRxiv&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;There are various preprint servers available out there (nicely listed &lt;a href=&quot;http://jabberwocky.weecology.org/2014/07/07/which-preprint-server-should-i-use/&quot;&gt;here&lt;/a&gt;)
and although the concept is the same for all,
the audience to which you want to reach out differs a lot.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;http://ergoso.me/general/2014/01/06/about.html&quot;&gt;As you know&lt;/a&gt;, I am studying Computational Biology, hence I am a fan of the new &lt;a href=&quot;http://www.biorxiv.org/&quot;&gt;bioRxiv&lt;/a&gt; service.
I have been subscribed to e-mail alerts for the subject areas that I am interested in since the service started.
I usually use my e-mail box as a task manager and label e-mails for adding them to my &lt;em&gt;have-a-look-at&lt;/em&gt; list.
This is especially useful for queuing papers of interest to me for reading them later.
And once I have the time, I go back to these marked e-mails, check these papers out by reading their abstracts and downloading the full article when necessary.&lt;/p&gt;

&lt;p&gt;I have also been doing the same thing for &lt;a href=&quot;http://www.biorxiv.org/&quot;&gt;bioRxiv&lt;/a&gt; e-alerts
and lately noticed something weird: &lt;strong&gt;some of my saved papers were not accessible on the site any more&lt;/strong&gt;.
The first time it happened to me, I thought this was a bug on the site and did not worry much about it.
The second time, however, was enough to get me suspicious.
So I went back to these missing papers and double checked the links on the &lt;a href=&quot;http://www.biorxiv.org/&quot;&gt;bioRxiv&lt;/a&gt;,
but all I was getting was an error on the site saying:&lt;/p&gt;

&lt;div class=&quot;quote&quot;&gt;
&lt;b&gt;Access Denied&lt;/b&gt;&lt;br /&gt;
You are not authorized to access this page.
&lt;/div&gt;

&lt;p&gt;Interesting enough, I found that these papers were not even being listed on the site!
Maybe, I thought, these papers were removed from the site on purpose, but is this even possible?
According to the documentation on the site, it is not supposed to happen:&lt;/p&gt;

&lt;div class=&quot;quote&quot;&gt;
... Authors may submit a revised version of an article to bioRxiv at any time and can update the bioRxiv record with a link to a version of an article that has been published in a journal. &lt;b&gt;Once posted on bioRxiv, articles are citable and therefore cannot be removed&lt;/b&gt;...
&lt;/div&gt;

&lt;p&gt;So I tweeted about it:&lt;/p&gt;

&lt;blockquote class=&quot;twitter-tweet&quot; lang=&quot;en&quot;&gt;&lt;p&gt;and when you go to the website for one of these, it says &amp;quot;access denied&amp;quot; -- a bug or a feature? More transparency on this would be great.&lt;/p&gt;&amp;mdash; B. Arman Aksoy (@armish) &lt;a href=&quot;https://twitter.com/armish/statuses/484495976494030848&quot;&gt;July 3, 2014&lt;/a&gt;&lt;/blockquote&gt;
&lt;script async=&quot;&quot; src=&quot;//platform.twitter.com/widgets.js&quot; charset=&quot;utf-8&quot;&gt;&lt;/script&gt;

&lt;p&gt;but got no response back.
Then I decided to see if I can scrape all such removed papers from the site to see if there are too may of these removed papers
and came up with &lt;a href=&quot;https://gist.github.com/armish/8693f1afbc80a7e6d503&quot;&gt;a really small script&lt;/a&gt; that records the HTTP response for successive biorXiv DOI numbers up to a point.
Based on this HTTP response, you can categorize the papers as follows:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;404&lt;/strong&gt;: the DOI has not been registered yet.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;200&lt;/strong&gt;: the DOI has been registered and the paper is available.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;403&lt;/strong&gt;: the DOI has been registered but access to the paper is restricted – &lt;em&gt;i.e.&lt;/em&gt; the paper has been removed.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Running the script and collecting this information for all DOIs from &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;10.1101/000001&lt;/code&gt; up to &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;10.1101/007000&lt;/code&gt;,
I found 498 published and 5 removed &lt;a href=&quot;http://www.biorxiv.org/&quot;&gt;bioRxiv&lt;/a&gt; papers.
&lt;strong&gt;This means that almost 1% of the &lt;a href=&quot;http://www.biorxiv.org/&quot;&gt;bioRxiv&lt;/a&gt; papers has been removed from the &lt;em&gt;archive&lt;/em&gt; with no explanation at all.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;It also turns out that although &lt;a href=&quot;http://www.biorxiv.org/&quot;&gt;bioRxiv&lt;/a&gt; has removed these articles, 
Google has not forgotten about them yet.
So for those that are curious, here is the list of removed papers and their titles:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;a href=&quot;http://dx.doi.org/10.1101/002055&quot;&gt;10.1101/002055&lt;/a&gt;: &lt;em&gt;Somatic mitochondrial DNA mutations are associated with progression, metastasis and death in oral squamous cell carcinoma&lt;/em&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;http://dx.doi.org/10.1101/002089&quot;&gt;10.1101/002089&lt;/a&gt;: &lt;em&gt;Protectome analysis: a new selective bioinformatics tool for bacterial vaccine candidate discovery&lt;/em&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;http://dx.doi.org/10.1101/002451&quot;&gt;10.1101/002451&lt;/a&gt;: &lt;em&gt;Genome-Wide Introgression Revealed Pervasive Hybrid Incompatibilities (HI) between Caenorhabditis species&lt;/em&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;http://dx.doi.org/10.1101/003251&quot;&gt;10.1101/003251&lt;/a&gt;: &lt;em&gt;SeqGL identifies context-dependent binding signals in genome-wide regulatory element maps&lt;/em&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;http://dx.doi.org/10.1101/005421&quot;&gt;10.1101/005421&lt;/a&gt;: &lt;em&gt;CRISPR/Cas9 nuclease-mediated gene knock-in in bovine pluripotent stem cells and embryos&lt;/em&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The number of removed papers might not be that big, but I think the situation is worrisome.
I am pretty sure there are valid reasons to why these papers were removed after they got posted on the site,
but I think if &lt;a href=&quot;http://www.biorxiv.org/&quot;&gt;bioRxiv&lt;/a&gt; is planning to be the &lt;em&gt;de facto&lt;/em&gt; preprint server for Biological Sciences,
then it should be either more strict about its terms of distribution or be more transparent about its decisions.&lt;/p&gt;

&lt;p&gt;What do you think?&lt;/p&gt;

</content>
 </entry>
 
 <entry>
   <title>TCGA Methylation Data - Part I: Motivation</title>
   <link href="https://ergoso.me/cancer/tcga/methylation/2014/02/26/methylation-1.html"/>
   <updated>2014-02-26T00:00:01+00:00</updated>
   <id>https://ergoso.me/cancer/tcga/methylation/2014/02/26/methylation-1</id>
   <content type="html">&lt;p&gt;As I &lt;a href=&quot;http://ergoso.me/metabolism/cancer/tcga/mutations/2014/01/08/muwheel.html&quot;&gt;mentioned earlier&lt;/a&gt;, the data available as part of the Cancer Genome Atlas (TCGA) project is incredibly useful for many types of analysis.
From a computational biologist point of view, a systematic analysis of all these data sets becomes interesting only when the analysis glues two or more different data types together (remember &lt;a href=&quot;http://ergoso.me/cancer/tcga/mutations/mgam/2014/01/28/mgam.html&quot;&gt;this post&lt;/a&gt;?).
People usually try to show that their analysis is strongly coupled to the survival data, 
meaning that their results might help explain how likely a patient is expected to survive in a given time frame, let’s say five years.
The survival data set is quite important for many reasons;
but people use it, mainly because it is often the only quantifiable phenotype that you can easily get out from the TCGA project.
This is also why you almost always see some sort of Kaplan-Meier plot as the last figure of computational biology papers that have something to do with cancer genomics.&lt;/p&gt;

&lt;p&gt;I am, of course, not critizing this approach – not at all.
I am simply saying this, because I’ve come to learn (quite late) that for such a computational approach to succeed, you need the following:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;Input data, &lt;em&gt;e.g.&lt;/em&gt; mutations or copy-number alterations&lt;/li&gt;
  &lt;li&gt;Quantifiable or categorical phenotype, &lt;em&gt;e.g.&lt;/em&gt; survival or subtype&lt;/li&gt;
  &lt;li&gt;A method that can use your input data to predict a phenotype with some success rate and also can tell you how it does it, &lt;em&gt;e.g.&lt;/em&gt; important features.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I am sure, by now, you realized that this is a typical machine learning problem and that is why using these algorithms are so important in this field.
As you can expect, there are many algorithms out there that can efficiently attack this problem;
but the actual problem is the availibility of the phenotype data.&lt;/p&gt;

&lt;p&gt;So the problem, actually, is that we don’t have much information about patients that can be used a direct phenotypic measure.
People have already realized that for different types of data,
there are ways of summarizing the multi-dimensional data in such a way that it can be used as a phenotype in an analysis like I mentioned above.
Not surprisingly, these efforts lead to many interesting biological associations, such as:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;Epigenetic &lt;em&gt;MLH1&lt;/em&gt; silencing and microsatellite instability (&lt;a href=&quot;http://hmg.oxfordjournals.org/content/8/4/661.short&quot;&gt;Simpkins et al., 1999&lt;/a&gt;)&lt;/li&gt;
  &lt;li&gt;&lt;em&gt;POLE&lt;/em&gt; mutations and extremely high mutation rates (&lt;a href=&quot;http://www.nature.com/ng/journal/v45/n2/abs/ng.2503.html&quot;&gt;Palles et al., 2013&lt;/a&gt;)&lt;/li&gt;
  &lt;li&gt;&lt;em&gt;IDH1/2&lt;/em&gt; mutations and hyper-methylation (&lt;a href=&quot;http://www.nature.com/nature/journal/v483/n7390/abs/nature10866.html&quot;&gt;Turcan et al., 2012&lt;/a&gt;)&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;As you know, I am really interested in characterizing possible outcomes of alterations in metabolic genes;
so after reading a bunch of papers about the &lt;em&gt;IDH&lt;/em&gt;-case and having lots and lots of discussions with &lt;a href=&quot;http://pseudon-ome.blogspot.com/&quot;&gt;Ed&lt;/a&gt;,
we decided to analyze the TCGA data to answer the following question:&lt;/p&gt;

&lt;div class=&quot;quote&quot;&gt;
	Do somatic alterations in any of the metabolic genes lead to an abbarent methylation profile in tumors, such as hypo- and hyper-methylation? 
&lt;/div&gt;

&lt;p&gt;At first, the question seemed to be easily approachable;
but to my surprise, it took us quite a lot of time to actually understand the nature of the data 
and devise ways to ask this question in a proper way.&lt;/p&gt;

&lt;p&gt;This was yet another pet project of ours and although it did not lead us to any interesting results, that is, results interesting enough to get published, 
I learned a lot from it and wanted to share all these things I get to learn in the process.
I will split everything into five different blog posts, each of them covering different aspects of the analysis, with first one being this very post as an introduction:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;Motivation&lt;/li&gt;
  &lt;li&gt;Hyper- and hypo-methylation as a phenotype&lt;/li&gt;
  &lt;li&gt;Learning with random forests and extracting important features&lt;/li&gt;
  &lt;li&gt;Genomic alterations that correlate with different methylation profiles&lt;/li&gt;
  &lt;li&gt;Wrap-up and more about methylation profiles&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;So stay tuned for new posts with some nice heatmaps, R scripts and interesting gene lists!&lt;/p&gt;
</content>
 </entry>
 
 <entry>
   <title>MGAM mutations in melanoma</title>
   <link href="https://ergoso.me/cancer/tcga/mutations/mgam/2014/01/28/mgam.html"/>
   <updated>2014-01-28T00:00:01+00:00</updated>
   <id>https://ergoso.me/cancer/tcga/mutations/mgam/2014/01/28/mgam</id>
   <content type="html">&lt;p&gt;Right after I came up with the &lt;a href=&quot;http://ergoso.me/metabolism/cancer/tcga/mutations/2014/01/08/muwheel.html&quot;&gt;μ-wheel&lt;/a&gt; visualization,
people immediately followed up with the obvious question: “&lt;em&gt;looks nice, but what is next?&lt;/em&gt;”&lt;/p&gt;

&lt;p&gt;If you, like I do, work in the field of Computational Biology, the chances are that you are already familiar with the following situation:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;Unbelievable amount of information is publicly available for analysis (e.g. &lt;a href=&quot;http://cancergenome.nih.gov/&quot;&gt;TCGA&lt;/a&gt;).&lt;/li&gt;
  &lt;li&gt;You pick one particular question, out of millions, and start crunching the data in a systematic manner to answer your question.&lt;/li&gt;
  &lt;li&gt;You end up with a huge list of results, &lt;em&gt;i.e.&lt;/em&gt; multiple answers to your question.&lt;/li&gt;
  &lt;li&gt;You browse these results and see that some of them are so good to be true, some not-so-bad and some not interesting to you at all.&lt;/li&gt;
  &lt;li&gt;And you start thinking about what to do with all these results.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I found myself clueless about what to do next (as step 6) and every time this happened, I started asking people around me for advice on what to do.
It took me four years of graduate-level education and hours (sometimes even days) of discussions to realize that I, often, do not have many options but to take either of the following two paths:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;Cherry-pick one really interesting result and try to go for a validation, that is often accompanied with wet-lab experiments.&lt;/li&gt;
  &lt;li&gt;Show that these results can be used, again in a systematic fashion, to predict a particular phenotype (e.g. survival time for each patient).&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;And I should mention that these paths are of interest to me only when I want to get some secondary yet &lt;em&gt;publishable&lt;/em&gt; results.
It is rather unfortunate that the primary results are often considered not to be sufficient for any type of publication,
hence they frequently get lost in the transition to secondary results due to various things (negative results, time, &lt;em&gt;etc.&lt;/em&gt;).&lt;/p&gt;

&lt;p&gt;Let’s get back to the two possible paths and try to focus on the former. 
The problem with follow-up experiments is that you need money, time, instruments and, of course, expertise on the subject of interest.
Computational Biology labs often lack one or more of these components, therefore have to look for collaborations, as it is supposed to be.
And here comes the problem of convincing somebody that what you want to do is interesting – and this problem is a really really hard one.&lt;/p&gt;

&lt;p&gt;How, in the world, you can convice your collaborator about doing experiments on that gene that nobody has ever heard of before but that also shines as a pearl in your result set?
Well… based on my experience, majority of time, you cannot – and I find this really sad.
&lt;em&gt;Really sad&lt;/em&gt; that we are just trashing some random observations like nothing just because they do not look promising.&lt;/p&gt;

&lt;p&gt;So that is why I am going to mention this gene, called &lt;em&gt;MGAM&lt;/em&gt;, in this post as an example to this situation, although I know that this is going to be a quite far-reaching post.&lt;/p&gt;

&lt;h3 id=&quot;mgam-maltase-glucoamylase-mutations-in-melanoma&quot;&gt;MGAM, maltase-glucoamylase, mutations in melanoma&lt;/h3&gt;
&lt;p&gt;While eye-balling the &lt;a href=&quot;http://ergoso.me/metabolism/cancer/tcga/mutations/2014/01/08/muwheel.html&quot;&gt;μ-wheel&lt;/a&gt;,
I realized that there is this gene called MGAM (maltase-glucoamylase) that have a few recurrent mutations across many cancer studies.
So I went to &lt;a href=&quot;http://cbioportal.org&quot;&gt;cBioPortal&lt;/a&gt;, as usual, and submitted a cross-cancer query to see how all mutations in this gene look like.
Below is a cross-cancer histogram showing relative frequencies of alteration (mutations) in MGAM:&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;http://bit.ly/LkstgY&quot;&gt;&lt;img src=&quot;/img/mgam-crosscancer-histogram.png&quot; alt=&quot;MGAM cross-cancer mutation histogram&quot; /&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Notice that the two studies to the far left, the ones that have the highest alteration frequency, are all &lt;em&gt;melanoma&lt;/em&gt; characterized by two different studies?
This is interesting – possibly indicating they do something in a tissue-specific way.
How about looking at how the mutations pile-up across all these studies:&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;http://bit.ly/LkstgY&quot;&gt;&lt;img src=&quot;/img/mgam-crosscancer-mutations.png&quot; alt=&quot;MGAM cross-cancer mutation pile-ups&quot; /&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Let’s assume that this gene is not one of the TTN-like &lt;a href=&quot;http://www.genome.gov/multimedia/slides/tcga1/tcga1_lawrence.pdf&quot;&gt;fishy genes&lt;/a&gt;.
Then, looks like, as was also apparent from the wheel, there are a few recurrent mutations in this gene.
Recurrent mutations, and the fact that the mutations have relatively high allele frequencies in mutated samples, are good signs:
it means that these mutations possibly help the cells in some way, hence they were not selected against and did not get lost in the process.&lt;/p&gt;

&lt;p&gt;Furthermore, most of these mutations seem to be &lt;em&gt;missense&lt;/em&gt; mutations (green) and only a few of them, &lt;em&gt;truncating&lt;/em&gt; mutations (red).
This, although weakly, implies the missense mutations are possibly &lt;em&gt;activating&lt;/em&gt; rather than &lt;em&gt;disabling&lt;/em&gt;,
because otherwise, we would expect to see relatively more truncating mutations.&lt;/p&gt;

&lt;p&gt;If activating, then what do all those mutations affect the function of the protein?
Considering that this gene is annotated as a &lt;em&gt;maltase-glucoamylase&lt;/em&gt; (ECs: &lt;a href=&quot;http://www.kegg.jp/dbget-bin/www_bget?ec:3.2.1.20&quot;&gt;3.2.1.20&lt;/a&gt; and &lt;a href=&quot;http://www.kegg.jp/dbget-bin/www_bget?ec:3.2.1.3&quot;&gt;3.2.1.3&lt;/a&gt;),
we can assume that for MGAM-mutated cells, the reaction this enzyme catalyzes runs more effectively and faster
– meaning that we have more breakdown of maltose-like species into glucose for cancer cells than it is for normal cells.&lt;/p&gt;

&lt;p&gt;This is interesting because we do know that cancer cells require more glucose uptake to support their incredibly high proliferation rate, again compared to the normal cells.
For the sake of the argument, let’s assume that the mutant protein causes more glucose availability hence is more advantageous to the cancer cells;
but what can we do about it?&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;http://bitbucket.org/armish/pihelper&quot;&gt;&lt;img src=&quot;/img/mgam-drugs.png&quot; alt=&quot;MGAM-targeting drugs&quot; /&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;No worries, I have some good news for you:
first, looking at the &lt;a href=&quot;http://bitbucket.org/armish/pihelper&quot;&gt;PiHelper&lt;/a&gt; data, looks like there are FDA-approved drugs (orange hexagons in the figure above) that inhibit the activitity of this gene and are already commercially available;
second, MGAM is a membrane protein with an extracellular domain, making it easy to raise potentially blocking antibodies against it.&lt;/p&gt;

&lt;p&gt;So, the over-reaching idea is as follows:&lt;/p&gt;

&lt;div class=&quot;quote&quot;&gt;
Recurrent mutations in MGAM gene are activating and this leads to more glucose availability 
and therefore is advantageous for cancer cells. Treating MGAM-mutated 
samples with one of the available targeted drugs to inhibit MGAM function, 
can help inhibiting proliferation of cancer cells to some degree. If so, developing antibody-based,
hence more specific, drugs against this protein is an option and it provides an 
opportunity towards a therapy alternative.
&lt;/div&gt;

&lt;p&gt;Then, the follow-up experiment to this hypothesis should be pretty straight-forward and relatively cheap:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;Obtain MGAM-mutated and MGAM-wildtype cell lines.&lt;/li&gt;
  &lt;li&gt;Obtain all available drugs targeting MGAM.&lt;/li&gt;
  &lt;li&gt;Test whether MGAM-mutated cell lines are more sensitive to inhbition with any of these drugs compared to MGAM-wildtype ones.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;But finding a collaborator or funding for these experiments?… 
Well, I don’t think it is going to be that easy.&lt;/p&gt;

</content>
 </entry>
 
 <entry>
   <title>μ-wheel</title>
   <link href="https://ergoso.me/metabolism/cancer/tcga/mutations/2014/01/08/muwheel.html"/>
   <updated>2014-01-08T00:00:01+00:00</updated>
   <id>https://ergoso.me/metabolism/cancer/tcga/mutations/2014/01/08/muwheel</id>
   <content type="html">&lt;p&gt;I am truly amazed by the abundance of the cancer genomics data and our current knowledge about the genes and the pathways they are involved in.
With all these computational-friendly data sets out there, 
it is not surprising that we see a lot of tools, algorithms, and pipelines that analyze the data sets in their own way and suggest &lt;em&gt;common&lt;/em&gt; mechanisms of tumor initiation or progression.
Thanks to these recent tools and really clever studies, we now have a somewhat satisfactory explanation for many cancer types and how they occur.&lt;/p&gt;

&lt;p&gt;Most of the explanations about tumor formation, however, focus on the genetic alterations of a limited set of genes.
These well-known genes are commonly called &lt;strong&gt;cancer genes&lt;/strong&gt; in the field 
and I will probably not be exaggerating if I say the community is obsessed with the cancer genes
– like &lt;em&gt;TP53&lt;/em&gt; or &lt;em&gt;EGFR&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;It is, of course, easier to study common events due to the sample size and the statistical power that comes with it;
but I believe, as a community, we are doing terrible when it comes to leverage all the knowledge we have accumulated over years about the common alterations to shed light on relatively &lt;strong&gt;rare alterations&lt;/strong&gt; that might contribute to tumor initiation.
There is a reason I am bringing this issue up: I believe there are ways we approach to this problem of rare alterations and exploring what their potential roles are.
The way to attack this problem should be no different than the widely-known schema of problem solving: &lt;em&gt;use prior knowledge to boost your statistical power to discover new associations&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;And in this post, I will try to explain one way of doing this – a project we (&lt;a href=&quot;http://pseudon-ome.blogspot.com/&quot;&gt;Ed&lt;/a&gt; and I) call &lt;strong&gt;μ-wheel&lt;/strong&gt;.
For this project, we decided to approach the cancer genomics data from the metabolism perspective, a point of view that is often under-estimated.
Sure, there are many well-studied areas in cancer metabolism, such as the &lt;a href=&quot;http://en.wikipedia.org/wiki/Warburg_effect&quot;&gt;Warburg Effect&lt;/a&gt;,
but we still lack a general understanding of how and if alterations in metabolic genes have an effect on the cell behaviour.&lt;/p&gt;

&lt;p&gt;To get an overview of alterations in metabolic genes, 
we decided to visualize the recurrent mutations in these genes in the hope for finding yet-another-hotspot
– something like the recently emerged &lt;a href=&quot;http://www.nature.com/nature/journal/v462/n7274/full/nature08617.html&quot;&gt;IDH story&lt;/a&gt;.
For this, we:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;Downloaded the cross-cancer mutation data from the &lt;a href=&quot;http://cancergenome.nih.gov/&quot;&gt;TCGA&lt;/a&gt; iniative (&lt;a href=&quot;http://dx.doi.org/10.7303/syn1710680.4&quot;&gt;doi:10.7303/syn1710680.4&lt;/a&gt;).&lt;/li&gt;
  &lt;li&gt;Using the &lt;a href=&quot;http://en.wikipedia.org/wiki/Enzyme_Commission_number&quot;&gt;Enzyme Commission Number&lt;/a&gt; &lt;a href=&quot;http://www.genome.jp/dbget-bin/www_bfind?enzyme&quot;&gt;database&lt;/a&gt;, classified all genes that have an EC number associated to it as &lt;em&gt;metabolic&lt;/em&gt;.&lt;/li&gt;
  &lt;li&gt;For all metabolic genes, extracted recurrent (n&amp;gt;1) mutations, that is mutations that are observed more than once and have a amino-acid-based coordinate to it.&lt;/li&gt;
  &lt;li&gt;Grouped all mutations in multiple levels in order to look for enrichment in any categorical level.&lt;/li&gt;
  &lt;li&gt;Created a web-based viewer, &lt;a href=&quot;http://ergoso.me/data/muwheel/&quot;&gt;μ-wheel&lt;/a&gt; for better exploration of the data (&lt;a href=&quot;http://dx.doi.org/10.6084/m9.figshare.894458&quot;&gt;doi:10.6084/m9.figshare.894458&lt;/a&gt;).&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href=&quot;http://ergoso.me/data/muwheel/&quot;&gt;&lt;img src=&quot;/img/muwheel-preview.png&quot; alt=&quot;Muwheel preview&quot; /&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;You can use the configuration options at the bottom of the page to customize the amount of information shown on the wheel.
The shade of the red represents how many recurrent mutations there are in a category and for inner circles, the number of recurrent mutations in all lower levels are summed-up.
Try moving your mouse over to different levels of rims and you will see that a pop-up will provide you with more information.
Clicking on the enzyme name will take you to a &lt;a href=&quot;http://www.genome.jp/dbget-bin/www_bfind?enzyme&quot;&gt;KEGG Enzyme&lt;/a&gt; page where you can learn more about the functions of those genes
and clicking on the gene name will take you to the &lt;a href=&quot;http://cbioportal.org&quot;&gt;cBioPortal&lt;/a&gt; where you can visualize the distribution of mutations in a better way, e.g.:&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;http://www.cbioportal.org/public-portal/cross_cancer.do?tab_index=tab_visualize&amp;amp;cancer_study_id=all&amp;amp;gene_list=IDH2&amp;amp;data_priority=1&amp;amp;_&amp;amp;Action=Submit&quot;&gt;&lt;img src=&quot;/img/idh2-muts.png&quot; alt=&quot;IDH2 mutations in cBioPortal&quot; /&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This visualization has some obvious problems:
First, some genes (such as &lt;em&gt;PTEN&lt;/em&gt;) show up in multiple EC categories, cluttering the view;
Second, the colors of the higher categories do not directly suggest enrichment in that category as these numbers are now normalized for the number of locations/genes under a specific category.
Third, some metabolic genes, in fact, are &lt;em&gt;cancer genes&lt;/em&gt; according to the &lt;a href=&quot;http://www.sanger.ac.uk/research/projects/cancergenome/census.html&quot;&gt;Sanger Cancer Gene Census&lt;/a&gt; (also have a look at &lt;a href=&quot;/img/metabolic-cancer-genes.png&quot;&gt;these charts&lt;/a&gt;).
These, especially the second one, are addressable problems, for example by applying a &lt;a href=&quot;http://www.broadinstitute.org/gsea/index.jsp&quot;&gt;GSEA&lt;/a&gt;-like enrichment analysis.&lt;/p&gt;

&lt;p&gt;Even in its current state, this visualization is really helpful in conveying the following message:
if we are to group rare and common alterations by some known categorical system (&lt;em&gt;prior knowledge&lt;/em&gt;), we might start to see enrichments that are not obvious when investigated one by one (&lt;em&gt;boosting power&lt;/em&gt;).
Try playing with the wheel, filter some mutation types out and change the recurrency levels; 
you will see interesting enzyme families starting to emerge from the whole mess.
Yet, be aware that before conducting a proper statistical enrichment, it is hard to claim anything from this analysis.&lt;/p&gt;

&lt;p&gt;But again, this is an on-going pet project and I am expecting to follow up on this soon.
So stay tuned and feel free to comment or get in contact;
I will love to hear your thoughts on this.&lt;/p&gt;
</content>
 </entry>
 
 <entry>
   <title>About</title>
   <link href="https://ergoso.me/general/2014/01/06/about.html"/>
   <updated>2014-01-06T00:00:01+00:00</updated>
   <id>https://ergoso.me/general/2014/01/06/about</id>
   <content type="html">&lt;p&gt;Hi there!&lt;/p&gt;

&lt;p&gt;This is a blog about various things, but it is mainly about &lt;em&gt;Computational Biology&lt;/em&gt;, &lt;em&gt;Systems Biology&lt;/em&gt; and &lt;em&gt;Cancer Research&lt;/em&gt;.
I started this blog, just because I want to share a few humble yet probably interesting observations–I had and hopefully will have–
and my experience throughout my research.&lt;/p&gt;

&lt;p&gt;As many of researchers already are, I am, too, frustated with the current system of publishing and making key results available for a larger communitee.
I, also like many others in the field, have been suffering a lot from the lack of published negative results.
I often find myself trying to investigate things that have already been studied extensively, yet there is not much information publicly available on it.
This is mainly because we do not bother publishing small findings if they are not &lt;em&gt;interesting&lt;/em&gt; enough to get published.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Ergosome&lt;/strong&gt; will hopefully become the medium that I will shortcut this painfully process and write things I find worthy of sharing.
The name of the blog is simply a synonym for &lt;a href=&quot;https://en.wikipedia.org/wiki/Polysome&quot;&gt;Polyribosome&lt;/a&gt;;
and it is also a word that resonates well with the famous Latin phrase “&lt;a href=&quot;http://en.wikipedia.org/wiki/Cogito_ergo_sum&quot;&gt;cogito, ergo sum&lt;/a&gt;.”&lt;/p&gt;

&lt;p&gt;Want to learn more? Feel free to visit &lt;a href=&quot;http://arman.aksoy.org&quot;&gt;my personal page&lt;/a&gt; or shoot me an e-mail at &lt;strong&gt;arman@aksoy.org&lt;/strong&gt;.&lt;/p&gt;

</content>
 </entry>
 

</feed>
