The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/AI & Data/I See Data People
I See Data People artwork

25 - The Judith Gu Episode

I See Data People · 2023-10-03 · 14 min

0:00--:--

Key moments - from our scoring

Substance score

50 / 100

Five dimensions, 20 points each

Insight Density11 / 20
Originality10 / 20
Guest Caliber14 / 20
Specificity & Evidence9 / 20
Conversational Craft6 / 20

Judy Gu brings deep experience from Morgan Stanley, Goldman Sachs, and Citadel to explain how modern equities quant trading operationalizes data at scale. Her team at Scotiabank runs systematic market-making across U.S. and Canadian equities, ingesting real-time tick data (quotes, trades, volumes, spreads) and building intraday technical signals to detect market dislocations and idiosyncratic stock risk. The evolution she's witnessed - from banks loading raw vendor data onto internal servers to consuming derived analytics via cloud APIs and streaming technologies - has materially shortened time-to-market. Gu emphasizes that low-latency data architecture splits into two workflows: reactionary (millisecond-level reactions to spread and volatility moves via convex optimization hedging) and predictive (alpha signals using in-memory time series and machine learning). Her most contrarian insight is that fundamental data - specifically news sentiment and event impact - should complement pure market data in short-term trading, a view she held a decade ago that is now gaining acceptance. She flags missing infrastructure: robust corporate action adjustments and historical security master tracking (to eliminate survivor bias in backtests) would unlock material efficiency gains. For modeling, she highlights Google's causal impact and synthetic control methods to isolate single-stock event impact, cautioning that model choice (excess returns vs. raw returns) shapes correlation structures and drives actionable insights. Looking ahead, she predicts large language models will project stock narratives and news sentiment, enabling longer-term positioning that reverse-engineers short-term tactics.

Key takeaways

  • →Real-time market-making requires dual-track data processing: reactionary signals at millisecond-to-microsecond latency for spread and volatility reactions, plus predictive alpha signals at minute-to-hour horizons using in-memory time series and machine learning.
  • →Corporate action time series and historical security master data remain critical infrastructure gaps that constrain backtesting quality and operational efficiency, despite being industry-wide pain points rather than novel content.
  • →News sentiment and fundamental data provide causal reasoning for short-term volatility and buy-sell imbalances that pure market data cannot capture, making integration of news analytics into market-making workflows increasingly valuable.
  • →Model architecture decisions - such as choosing between excess returns and raw returns as inputs - fundamentally shape correlation structures and the actionability of derived insights, not just data content.
  • →Large language models predicting future stock narratives and news sentiment (rather than just reacting to current events) will enable longer-term projections that can improve short-term tactical execution.

Guests

Judy Gu

Topics in this episode

Scotiabank U.S. Equities Sales and TradingReal-time tick data and MBBO (market best bid and offer)Convex optimization for hedgingIntraday technical signalsFactor risk modelsSynthetic control modelsCausal impact (Google)News sentiment analysisCorporate actions (dividends, splits, delistings)Security master data

Questions this episode answers

How do market makers use real-time tick data to manage risk on a millisecond timescale?

Market makers like Scotiabank's equities desk react to tick-by-tick changes in spreads and volatility using a pricing engine that adjusts positions on millisecond or microsecond latency, while continuously running convex optimization to rebalance hedges against their own positional risk changes.

What is the difference between reactionary and predictive signals in quantitative market making?

Reactionary signals (milliseconds to microseconds) respond to immediate market conditions like spread moves and volatility surges, while predictive alpha signals (minutes to hours) use stateful time series data and machine learning to project near-term price movements and relative return trajectories based on technical, news, and liquidity conditions.

Why does Judy Gu think news sentiment data should be part of market-making, not just market data?

News sentiment provides causal reasoning for price, volatility, and order imbalance changes that pure market data alone cannot explain, adding a forward-looking dimension that helps capture single-stock idiosyncratic risk and near-term event impact.

What data infrastructure does Scotiabank currently lack that would improve backtesting and operations?

Corporate action time series adjustments (dividends, splits, delistings) and historical security master tracking (to follow ticker changes and eliminating survivor bias) that refresh daily across the full historical dataset would materially reduce operational burden and backtest quality risks.

How are synthetic control and causal impact models being applied to news event analysis in market making?

These models quantify the isolated impact of news events on single stocks' relative performance by comparing the stock's behavior post-event to a synthetic control of similar stocks, requiring careful choices like using excess returns to maintain pre-event correlation for valid counterfactual estimation.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

11 / 20

The episode packs real technical content into 14 minutes - synthetic controls for news event impact, the reactionary vs. alpha-signal framework, and LLM-based narrative projection are substantive. However, the value is heavily siloed in quant finance and many ideas are explained at a conceptual level without enough follow-through for a practitioner to act on.

we are looking into using modeling techniques called synthetic controls and synthetic intervention, or a more recent model from Google's causal impact machine learning models to quantify near-term event impact on stocks' relative performance post-news event
access return by definition reduced the correlation across the stocks, even though the stocks can be from the same sectors

Originality

10 / 20

The idea of using causal-impact models for single-stock news events in a market-making context is genuinely non-standard, and predicting future narratives via LLMs is an interesting framing. But the guest's own 'controversial' view - that fundamental data matters in short-term trading - is by her own admission now mainstream, blunting the originality claim.

news sentiment was once considered as an alternative data is now becoming more mainstream. And my novel idea was some controversial view I had actually from a decade ago. is becoming more palatable today
we don't construct portfolios and allocate risks like buy-side investors. But rather, our risk is predominantly decided by our client order flows

Guest Caliber

14 / 20

Judith Gu is a legitimate practitioner who built a quant trading business at Scotiabank after VP-level roles at Goldman Sachs and Citadel - she has genuinely done the thing at scale. The short runtime limits how much depth is extracted from her credentials, but the credential-to-content ratio is solid.

my team runs the equities-quant trading on a market-making desk. At Scotiabank, we cover both Canada and U.S.
Over five years ago, when we started building this quant trading business, we never had to load any raw data from vendor

Specificity & Evidence

9 / 20

The episode names specific techniques (synthetic controls, Google's causal impact model, vector databases) and specific data artifacts (corporate action adjustments, security master, NBBO), which is creditable. However, there are zero performance figures, no concrete outcomes from implementing any approach, and no timelines beyond 'over five years ago' - the specificity is terminological rather than evidential.

Corporate actions like cash, stock dividend, stock splits happens every day given the stock universe. And historical prices and returns need to be adjusted to get the correct returns
The collaboration between cloud technology and other database, more performance-driven database like Vector Database, can bring down the technology barriers in a material way in the future

Conversational Craft

6 / 20

The hosts use a rigid, pre-planned question template ('what data do you wish you had?', 'most powerful insight?', 'most controversial view?', 'where is data going in five years?') with no meaningful follow-up when interesting threads appear, and they respond to substantive answers with flattery rather than probing. The interview functions as a structured promotional feature, not a real intellectual conversation.

Sounds like your controversial opinions are pretty prescient about the future
Love this. And you know quantitative trading often requires also low latency data solutions

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Most-used words

data46market23stock17risk16term14news13model12level10short10changes9judy8trading8making8real8stocks8intraday6

Episode notes

Judith is a Managing Director and Head Equities Quantitative Strategist for Scotiabank’s US Equities Sales and Trading. Prior to Scotiabank, Judith was a VP at Goldman Sachs and Citadel, after starting her career at Morgan Stanley. Judith discusses how market makers leverage data to hedge trading risk. She delves into why market makers increasingly need to leverage fundamental data, in addition to traditional market data, to effectively hedge risk. Judith also alludes to the market making use cases for alternative data and explains why trading operations need to have complete control over their technology infrastructure, particularly regarding trade execution.

Full transcript

14 min

Transcribed and scored by The B2B Podcast Index.

Hi, and welcome to another episode of IC Data People. We're joined today by Judy Gu. Judy is a Managing Director and Head Equities Quantitative Strategist for Scotiabank's U.S.

Equities Sales and Trading. Prior to Scotiabank, Judy was a VP at Goldman Sachs and Citadel after starting her career at Morgan Stanley. Judy, welcome to IC Data People. Thank you, Evan and Omri.

It's great to be here. It's fantastic to have you here, Judy. And you developed an enormous amount of expertise in quantitative trading. I'd love for you to tell us a little bit about how you utilize data on a day-to-day basis and how that data usage maybe has changed over the course of your career.

Oh, yeah, absolutely. So my team runs the equities-quant trading on a market-making desk. At Scotiabank, we cover both Canada and U.S.

And as a systematic market maker, we continuously post bid and offers for block size and auto-matching client's respondent orders based on our bids and offers. And then we're managing capital risk through a continuous hedging optimization and transaction algo monitoring. So our brand and butter data set is the voluminous real-time tick data streaming, including quotes, trades, volume, intraday volatility, spread changes on the stock level. And we also run a convex optimization using daily factor risk model.

So aside from this bread and butter data, we derive our own intraday technical signals to gauge market dislocations and stock idiosyncratic risk. We're in the process of researching intraday news event sentiments impact on stock's short-term returns. So the changes I've seen over the years in the market making space is the increasing sophistication of real-time analytics from both data content and technological advancements. So instead of just knowing the last data points, like the last trade, the last bid, the last MBBO, we can do analytics holding a series of time series data in memory.

So we can do a trend analysis and mean reversion analysis. So in addition, building a real-time multi-factor model, looking at fundamental updates like news, technical signals like real-time dislocations, liquidity factors, all commingled into guiding our risk management is another big trend developed over the years, over recent years. In addition to that, I'd also like to mention the transformative changes of how clients like ourselves receiving data over the years. And not very long ago, banks depend on data technology groups to load all the raw data feeds from vendors onto their own database servers.

The process can be challenging and costly. And this process has been transformed towards efficiency in both time and cost through cloud technologies and streaming technologies like messaging queues and APIs. Over five years ago, when we started building this quant trading business, we never had to load any raw data from vendor, but rather use them to immediately derive our own analytics data, which shortens the time to market and hitting our commercial lines in a more efficient manner.

Love this. And you know quantitative trading often requires also low latency data solutions And it becomes more important with time I believe And can you explain how timeliness factors into your workflow into your models into your day-to-day life? Oh, absolutely. That's a great question.

You're spot on, Omri. So systematic market making and risk management is all about real-time decision making. They're either very short term, such as milliseconds or microseconds, or quite short term, which can last from minutes to hours. So mainly there are two different categories.

The first one is reactionary. For example, the spread moves and volatility surges on tick-by-tick level, our market-making pricing engine reacts to stock-level market conditions on that milliseconds or microseconds level. In addition to reacting to quotes, we also react to our own positional risk changes on a continuous basis because our risk profile can change on a dime. From sector risk from different sector risk rotations to long short changes over the intraday courses, it can change on time.

So we monitor that as well. As a market maker, our risk exposure can change on time, as I said, and we run a convex optimization in real-time basis to adjust hedges. The second category is less reactionary, but carries some short-term predictive value. Those are what we call alpha signals, which try to make projections of near-term price movements or relative return trajectory, given the current condition, including technical news and liquidity changes.

Alpha signals requires stateful data, such as in-memory time series and prior conditions, such as recent stock sentiment and events. So those data points will either be calibrated through statistical or machine learning models on historical basis for the intraday parameters or calculated through intraday time series to make near-term predictions. So both reactionary and near-term predictive trajectory signals all happens in a very low latency space. Thanks so much for this.

And one very interesting point that we usually don't get from the show is what types of data set you wish you had access to, but you currently don't have? Yeah, that's a great question too. It talks about my daily pain, the things that I wish were there already. So it might be surprising to most people that the data we wish to have is actually more infrastructure-oriented data rather than content-focused data in our space.

So there are two types of data that I wish I had. One is the one that examples including very solid corporate action time series adjustments become available at fingertip. So corporate actions like cash, stock dividend, stock splits happens every day given the stock universe. And historical prices and returns need to be adjusted to get the correct returns.

One of the tricky parts of corporate action adjustment is that they need to be refreshed on a daily basis for the entire historical history. And to make such data available at will be very helpful to both market makers and buy side. Same comments goes to historical security master updates. Companies can change their tickers, go from public to private and then back to public again delisting spin and divestitures To be able to track them through history is critical for backtesting quality such as eliminating survivor biases for stocks that went bankrupt and lost track So those two, like corporate actions and the security master historical adjustments will be very critical and helpful to us if it can be available at our fingertip.

And yet another data that I wish I have is some of the derived analytics data without the need for us to ingest all the raw data to construct ourselves. For example, at our risk management level, we don't need all the level two and level three market data. However, there's some key liquidity insights requires level two and level three data to become complete. So rather than for us to ingest all the raw data and come up with those insights, it will be a lot more efficient for us if we are able to get those derived metrics directly.

So Judy, you've touched on a number of data sets that you wish you had access to. I'm curious, from what you do have access to or what you've seen recently, what's the most powerful insight that you've seen derived from data in recent memory? Yeah, sure. Great question.

So some of the interesting, if not powerful insights that are unique to market making is actually called the single stock risk domination. So as market makers, we don't construct portfolios and allocate risks like buy-side investors. But rather, our risk is predominantly decided by our client order flows. So single stock idiosyncratic risk and dislocations are very frequent in our space.

This phenomenon is nothing new. It's not a new discovery, but it still remains as a core risk management to market makers that needs ongoing research and adaptation. I believe the key to target this problem is going through the right type of alpha research and calibrations to target as those idiosyncratic risk at stock level. So, for example, to use a new sentiment research process as an example, we're not doing a portfolio-based cross-sectional research that constructs top docile and bottom docile for long-short strategies.

That's not our business model. We don't do that. But rather, we calibrate different news events' near-term impact on single stocks so we can capture a single stock's idiosyncratic move. And to model that, we are looking into using modeling techniques called synthetic controls and synthetic intervention, or a more recent model from Google's causal impact machine learning models to quantify near-term event impact on stocks' relative performance post-news event.

Here, I'd like to add that we can't just talk about data in absence of modeling decisions. Majority of the time, we use a stock's access return as inputs to models, access return being the return residuals after applying a factor model. However, one of the model requirements for synthetic control is the high correlation across stocks in the pre-intervention, like before the news happened. So access return by definition reduced the correlation across the stocks, even though the stocks can be from the same sectors.

But on the other hand, using stock returns or stock prices may introduce confounding factors like sector or market move. The point here is it's not just about data content, but also about the model. It how its model generates insights that drives the actionables Judy I curious what is your most controversial Yeah that a fun question So one of my more controversial views and beliefs is to adopt and incorporate fundamental data to market-making ecosystem. The conventional industry consensus is that we can derive most, if not all, market intelligence from market data alone in a very short-term trading like market making.

So the importance of the market data is undeniable and will always be there. Will always continue to be impactful in our business. But on the other hand, having worked on a trading floor for years, I recognized the impact of real-time fundamental updates like news event sentiment to short-term trading behavior. So news, in addition to market data, like spread changes, volatility changes, news provides a dimension that pure market data doesn't, which is it provides a causal reasoning of those changes.

The causal reasoning for short-term returns, volatilities, and buy-sell imbalances. So news sentiment was once considered as an alternative data is now becoming more mainstream. And my novel idea was some controversial view I had actually from a decade ago. is becoming more palatable today.

Sounds like your controversial opinions are pretty prescient about the future. So where do you see the data world going in the next five years? What is going to be popular five years from now or commonplace five years from now that maybe people aren't thinking about today? Yeah, I would say, Evan, five years is a very long time horizon to predict.

But I do see two growing trends worth following up. One is in the large language model development. I have read and heard some of the research trying to project upcoming narratives of the stock, the upcoming narratives or sentiments, even news using the large language model conditioned upon the current stocks, fundamentals and sentiment. So it's almost like predicting what's going to happen to the stock in those spaces, even predicting news for the stocks using a large language model.

So this is not the same as predicting return impact given the current condition. The research is trying to project the next likely news and market sentiment. And if we can get that right, we can use it for our longer-term projection and then reverse engineering out what we can do better in the short term. And from technological development, I see more combined power of handling large data sets for real-time analysis.

It's always a game about handling large and fast data, or sometimes people call that vast data, B-A-S-T, vast data, as in volume and speed. The collaboration between cloud technology and other database, more performance-driven database like Vector Database, can bring down the technology barriers in a material way in the future. Judy, that's, I think, an ambitious prediction about the future and what I'm excited to see. given your prescience in the past, I suspect you're probably right.

But most importantly, thank you so much for being a guest on IC Data People. It's great to be here. Thanks for inviting me.

More from I See Data People

All episodes →
  • 24 - The Richard Hoffmann Episode
  • 23 - The Jeremy Baksht Episode
  • 22 - The David Rosen Episode
  • 21 - The Kirk McKeown Episode
  • 20 - The Second Recap Episode
Explore the best B2B AI & Data podcasts →
All I See Data People episodes →