<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Lab Automation | Shunyang Wang</title><link>https://shunyang.xyz/tag/lab-automation/</link><atom:link href="https://shunyang.xyz/tag/lab-automation/index.xml" rel="self" type="application/rss+xml"/><description>Lab Automation</description><generator>Wowchemy (https://wowchemy.com)</generator><language>en-us</language><lastBuildDate>Sun, 07 Jun 2026 00:00:00 +0000</lastBuildDate><image><url>https://shunyang.xyz/media/icon_hua2ec155b4296a9c9791d015323e16eb5_11927_512x512_fill_lanczos_center_3.png</url><title>Lab Automation</title><link>https://shunyang.xyz/tag/lab-automation/</link></image><item><title>How to Automate the Last Mile in the Lab</title><link>https://shunyang.xyz/posts/automating-the-last-mile-in-the-lab/</link><pubDate>Sun, 07 Jun 2026 00:00:00 +0000</pubDate><guid>https://shunyang.xyz/posts/automating-the-last-mile-in-the-lab/</guid><description>&lt;p>Orchestration, lab informatics systems such as LIMS and ELNs, liquid handlers, and robotic arms are hot topics in lab automation. Recent examples include &lt;a href="https://doi.org/10.1038/s41586-023-06792-0" target="_blank" rel="noopener">Coscientist&lt;/a>, which connected a large language model with laboratory tools, and &lt;a href="https://doi.org/10.1038/s41586-023-06734-w" target="_blank" rel="noopener">A-Lab&lt;/a>, which connected machine learning, robotics, and material characterization. OpenAI and Ginkgo Bioworks also recently used GPT-5 with a cloud lab to run closed-loop experiments for &lt;a href="https://openai.com/index/gpt-5-lowers-protein-synthesis-cost/" target="_blank" rel="noopener">cell-free protein synthesis&lt;/a>.&lt;/p>
&lt;p>These systems can automate experiments and speed up the testing and validation of ideas. However, they can also be expensive and time-consuming to build. Today, they are often better at replacing repeated operations than replacing the daily scientific decisions made by scientists. Even in the OpenAI and Ginkgo study, human oversight was still needed for protocol improvements and reagent handling.&lt;/p>
&lt;h2 id="what-is-the-last-mile-in-the-lab">What is the last mile in the lab?&lt;/h2>
&lt;p>There are also bottlenecks where a small investment can save a lot of valuable scientists&amp;rsquo; time, but a general solution may not fit. These projects may not support a new business by themselves. They often need highly specific optimization and close collaboration with scientists. I call this the last mile in lab automation: the small but essential steps between a working tool and an end-to-end scientific workflow.&lt;/p>
&lt;p>For example, identifying the structures of impurities or drug metabolites often requires evidence from LC-MS, MS/MS, NMR, and other experiments. MS alone may not identify the exact position of a modification or distinguish isomers, so scientists need to cross-check different data and do a large amount of manual annotation (&lt;a href="https://pubmed.ncbi.nlm.nih.gov/17405144/" target="_blank" rel="noopener">Prakash et al., 2007&lt;/a>).&lt;/p>
&lt;p>Another example is choosing liquid chromatography methods for purification. This can require method screening, analytical runs, and then preparative runs. Each round needs decisions and setup and can use a lot of samples, solvents, and other consumables. Published workflows also describe screening analytical conditions before scaling promising methods to preparative purification (&lt;a href="https://pubmed.ncbi.nlm.nih.gov/21122868/" target="_blank" rel="noopener">Font et al., 2011&lt;/a>). What if we can automate some of these experience- and knowledge-based decisions?&lt;/p>
&lt;h2 id="why-the-last-mile-remains-manual">Why the last mile remains manual&lt;/h2>
&lt;p>These last-mile tasks often remain manual because the work is spread across different instruments, file formats, and software. The person who understands the scientific question may not own the instruments or the data systems. The workflow may also change from one project to another and still needs validation.&lt;/p>
&lt;p>In the structure annotation example, evidence from different instruments needs to be connected before a scientist can make a decision. In the chromatography example, results from one screening run affect the setup of the next run. The difficult part is not always running the instrument. It is connecting the data, decisions, and next actions.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img src="./last-mile-workflow.jpg" alt="A last-mile automation loop connecting a scientific question, an instrument, scattered data, a scientist&amp;amp;rsquo;s decision, and the next experiment." loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>&lt;em>The last mile connects the scientific need, laboratory work, data, decisions, and the next experiment.&lt;/em>&lt;/p>
&lt;h2 id="find-the-real-bottleneck">Find the real bottleneck&lt;/h2>
&lt;p>Finding the real bottleneck is why we automate the last mile. We are here to help scientists reduce the problems that take the most time and effort, not to force a top-down automation plan across the whole lab. This also gives us the advantage of working closely with scientists, understanding their real needs, and getting them fully involved in the project.&lt;/p>
&lt;h2 id="learnings">Learnings&lt;/h2>
&lt;p>Through my work, I have had the advantage of working closely with scientists from different parts of the drug discovery life cycle and with very different backgrounds. These are the lessons that I found general enough across projects.&lt;/p>
&lt;h3 id="you-need-to-win-the-trust">You need to win the trust&lt;/h3>
&lt;p>We are in a new period of automation, and no one has all the answers. The old methods are already validated, robust, and familiar to scientists. We need to get scientists on board by winning their trust.&lt;/p>
&lt;p>We should start small and always provide verifiable results. This does not mean that we cannot talk about the big picture, but we should plan in phases. When we reach each milestone, we gain more trust and move closer to the big picture. Do not think that starting small is a waste of talent. During this process, we can build the infrastructure, identify barriers, and change direction more easily. We can also show what we are capable of and where the limits of automation are. This will help with future collaborations.&lt;/p>
&lt;h3 id="get-a-prototype-as-soon-as-possible">Get a prototype as soon as possible&lt;/h3>
&lt;p>You and your collaborators may come from very different backgrounds. Sometimes, after a long discussion, both sides still understand the problem differently. It can feel like you are talking past each other. In this case, an early prototype can help communication.&lt;/p>
&lt;ol>
&lt;li>Scientists may not be fully clear about what they need. A hands-on demo can help them clarify the request.&lt;/li>
&lt;li>They will have a better understanding of the complexity.&lt;/li>
&lt;li>A quick prototype can prevent you from building too much before collaborators ask to start over.&lt;/li>
&lt;/ol>
&lt;p>This is a bitter lesson I have learned more than once.&lt;/p>
&lt;h3 id="save-as-much-useful-data-as-possible">Save as much useful data as possible&lt;/h3>
&lt;p>As a machine learning scientist working in lab automation, one of my most common and time-consuming problems is data. Why do we not collect data intentionally while building new tools?&lt;/p>
&lt;p>A mass spectrometry database built for spectral matching can also save method information and retention times, even if they are not needed at first. These data may later support retention time prediction, while the method information can help define how transferable a model is.&lt;/p>
&lt;p>We should also help guide data collection and point out missing data that may be useful later. For example, if a model shows that changes in the environment over time are more important than a few separate readings, we can start collecting more time points for cell culture.&lt;/p>
&lt;h3 id="separate-models-from-complex-tasks">Separate models from complex tasks&lt;/h3>
&lt;p>I will discuss this more in another article. Briefly, we should build the system in modules and define the scope of the current model. What are its inputs and outputs? How does it fit into future steps? Saying no to a request because it is outside the current scope can make the discussion much easier.&lt;/p>
&lt;h3 id="leave-extendable-plugins-for-the-future">Leave extendable plugins for the future&lt;/h3>
&lt;p>Based on the last lesson, we should also think about future extensions. This includes plugin design, database schemas, model transferability, and new modalities. For example, a workflow built for small molecules may reuse the same pipelines and experiment setup for peptides while changing the machine learning model in the backend.&lt;/p>
&lt;h2 id="closing-thoughts">Closing thoughts&lt;/h2>
&lt;p>The last mile may not look as exciting as a fully automated lab. However, these small and highly specific improvements can remove real bottlenecks from scientists&amp;rsquo; daily work. Large companies and startups may first capture the low-hanging fruit and the most profitable parts of lab automation. The real challenge of last-mile automation starts after that. These problems are highly specific to each organization and depend heavily on internal collaboration and development. This is why the practical approach described above matters: start from scientists&amp;rsquo; needs, keep the big picture in mind, and move toward it in small steps. Small results help build trust, prototypes help us learn, and useful data supports the next steps.&lt;/p></description></item></channel></rss>