
Industrial AI Podcast · 2026-07-01 · 50 min
Key moments - from our scoring
Substance score
44 / 100
Five dimensions, 20 points each
TiRex-2 is the successor to the original TiRex foundation model developed by NXAI and Johannes Kepler University Linz, moving from univariate to multivariate time series forecasting in a zero-shot regime. Levente Zoliomi, a PhD researcher at NXAI under Professor Sepp Hochreiter, explains that the core advancement enables the model to process multiple related time series simultaneously - called covariates or multivariate data - which are ubiquitous in industrial settings but far more complex than single signals. The model is built on a recurrent XLSTM architecture (a modernization of LSTM), which offers significant computational and memory advantages over transformer-based competitors like Amazon's Kronos II. TiRex-2 distinguishes between past covariates (observed up to the forecast point, like rainfall affecting traffic) and future covariates (known in advance, like holiday dates affecting retail sales), while maintaining strict causality to prevent future information leakage. This episode is essential for operations leaders evaluating time series foundation models, manufacturing engineers working with sensor data, and anyone deploying forecasting across supply chain, infrastructure, or financial applications who need to understand multivariate capabilities and architectural trade-offs.
TiRex-2 extends the original TiRex model to handle multivariate time series data through covariates, allowing it to process multiple related time series simultaneously rather than only a single variable, plus additional architectural improvements to the recurrent XLSTM base.
Past covariates are additional time series observed up until the forecast point (like current rainfall affecting traffic prediction), while future covariates are known in advance (like holiday dates driving retail sales), and TiRex-2 maintains strict causality to prevent future information leakage into past predictions.
XLSTM processes time series sequentially with constant memory size, requiring only linear computational and storage costs, whereas transformers store and compare all timesteps, resulting in quadratic memory and processing costs that become prohibitive in deployment scenarios.
Zero-shot means the model forecasts unseen data without prior training on it, which is critical for proprietary industrial sensors; however, NXAI offers optional fine-tuning to focus the model on specific patterns and frequencies relevant to a company's particular machinery and sensor data.
While users must define candidate covariates as domain experts, TiRex-2 inherently learns which ones meaningfully contribute to the forecast and applies different emphasis weights, though providing thousands of variables is unrealistic.
Our reviewer’s read on each dimension, with quotes from the episode.
There are genuine technical nuggets - constant vs. quadratic memory, streaming update mechanics, parameter count trade-offs, and past vs. future covariate causality - but they are buried under extended definitional throat-clearing on basics like 'what is a time series' and 'what is zero-shot.' The insight-to-filler ratio is mediocre for a practitioner audience.
the multivariate model of TireX-II is eighty two million parameters And Chronos II has one hundred twenty millions. so there's I would say an advantage of Tyrex too. But what's also something unique is if somebody decides that they do not need the multivariate capabilities of Tyres, you can turn off those sets of parameters and then this just gets reduced to thirty eight million
whenever we have now a new observation come in... We can just do a single quick step of, okay now we put this new information into our fixed memory enabled by this recurrent architecture and you just do another forecast. Meanwhile on something like chronos two or any other transformer based model in general that is not fully causal he would have to recompute A lot more off the past
The architectural argument for recurrent xLSTM over transformers in time series (constant memory, sequential processing, streaming efficiency) is a genuine differentiator worth hearing, but the broader framing - foundation models for time series, zero-shot forecasting, covariate handling - mirrors standard discourse in the field with no contrarian or first-principles challenge to dominant assumptions.
even if you were To compute for a million time steps You will have the exact same megabytes Of memory occupied on your system If you compare Just two ten times steps whereas with a transformer very, very quickly into something that's not really manageable
this strict causality especially in the future is important because it of course In time series forecasting. we cannot let any future information leak back into The past as that will just make the target too easy for the model
Levente is a genuine technical practitioner building the actual product discussed, with clear command of the architecture, but he is approximately one year into a PhD with no track record of deploying systems at scale; he is a researcher, not an operator or senior practitioner who has shipped production ML to industry customers.
I started my PhD about a year or so ago here in Linz and then also been at NXAI for about the same time as well
I am part of research team also PhD students my day much a blend of what a PhD student would do, reading papers catching up on new research and then also gathering ideas
The episode offers a handful of concrete numbers (82M vs 120M parameters, 38M in univariate mode, sub-100MB model size, GiftEval and FevBench benchmark names) but provides no actual benchmark scores, no customer case studies, no latency figures, and vague claims like 'ranking very much on top' that a decision-maker cannot act on.
the multivariate model of TireX-II is eighty two million parameters And Chronos II has one hundred twenty millions
you can shrink this down to, i would say under a hundred megabytes for something like Tirex
The host spends the majority of the episode asking entry-level definitional questions ('what is time series data,' 'what is a foundation model,' 'what is zero-shot') appropriate for a general-public explainer, not a B2B decision-maker podcast; there is no pushback, no challenging of benchmark claims, and the host's personal anecdotes frequently interrupt and slow the conversation.
this time series foundation model there those are two important things again at the majority of the listeners will know about remind us what our timeseries data? What are time series data used for
Is That The Standard Operating Mode? Can I at all.
Computed from the transcript - who did the talking, and the words that came up most.
Discover how TiRex-2 is revolutionizing time series AI for industry. Multivariate, streaming, and more - get the inside story. In this episode, I sit down with Levente Zolyomi, a leading PhD researcher at NXAI, to unpack the next evolution in time series AI: TiRex-2. We explore how TiRex-2 builds on its predecessor by handling complex multivariate data and streaming scenarios, opening new frontiers for industrial forecasting. Levente shares the story behind TiRex-2’s architecture, its breakthrough capabilities, and what sets it apart from transformer-based models like Kronos. I ask the questions you’re thinking - about zero-shot forecasting, synthetic data, and real-world benchmarks - so you can understand what matters most for your business. If you’re navigating industrial AI or just curious about the future of time series models, this conversation is packed with the insights you need.
Transcribed and scored by The B2B Podcast Index.
This podcast is presented by NXAI, your partner for time series foundation models and physical AI. Hi there! Welcome to a new episode of the Industrial AI Podcast. My name is Peter Sieberg And I'm you host.
Today i am going be talking to Levente Zoliomi. I hope I pronounced that correctly. We'll hear in moment. Levente is PhD researcher at NXAI He & I.
Today, I'm going to be talking about Tyrex too. More specifically about generalizing TyreX to multivariate data and streaming. Hi Leventi! Hello Peter nice to be here happy to talk about these topics.
exactly Nice for me as well knowing that you are the specialist on exactly this very topic. Let's start Levent by you introducing yourself to our listeners, please. Yes of course! So as she said I am a PhD researcher at NXAI and i'm also doing my Phd at the Johannes Keppler University here in Linz Austria under Professor Staphochreiter.
we have this nice time series research group that Many of my other brilliant PhD colleagues, we were working on Tyrex too. And I started my PhD about a year or so ago here in Linz and then also been at NXAI for about the same time as well. Okay We may be talking about that to very end but how long does your PhD take typically? Or you know How long is it gonna Take further until are going to ready finished I?
that depends on, of course a lot of circumstances. Usually between three to five years is the standard here? Very good one more maybe more personal question where does your name come from? has it got to do anything with?
It does not have to do anything with the wind. You know why I asked that, because i was thinking Levante is a word and name of specific winds coming from where? So if it doesn't come then where does it originate from ? so I grew up in Hungary And this actually very traditional Hungarian name Levente It has nothing to do with the wind.
Okay, it is a coincidence and there's only one letter difference between the spelling of my name. And okay you know that quite common. very good Let's start with tyrex two. Please give us a very quick in overview introduction and then we'll move into The separate details.
Yes, of course more than happy to. So Titex-II is the successor for the original model also developed by NXAI and the JKU. And the main pitch for Titext-II Is that it's still a time series foundation model. It is in this zero shot forecasting regime.
But now... The main improvement is that it can now handle multivariate data, and these covariates are actually abundant everywhere in industrial applications but a lot more complicated to process than just single signal. And the main idea behind IREX-II was expand this multivariat regime... and also we built quite few nice improvements on top of this already existing strong base.
That's the high-level overview, very good to over view. you gave me about four or five keywords that I'm just going to be asking you about. we're gonna be stepping back a little bit so We all have the same baseline and will take it from there. So first question is Tyrex too right?
Is the follow up off tyrex What was Tyrax, what is still today? Tyraxx? and maybe as we were talking like age how long does Tyraax exist. Yes of course let's go into that then.
so TyraX the original Tyra X model was released in twenty-twenty five around May And this one of first time series foundation models and especially one of the first ones that moved away from this transformer architecture, which is so prevalent everywhere. And later on we can get into detail why it's nice to be using something more efficient than a transformer. but in a nutshell Maybe we talk or is that now? if as were talking a little bit to the past, but I'll come to it in a minute anyway.
Second question then this time series foundation model there those are two important things again at the majority of the listeners will know about remind us what our timeseries data? What are time series data used for and water? some different use cases mean. Of course, we're talking here.
We are the industrial AI podcast. assume that is one use case right? But there's completely other use cases as well. so maybe you talk a little bit about two three specific areas.
Yes of course. So time series data in many places In fact I would say more place than most people will realize. Of course industry is one. to give maybe a bit more specific examples, for example in manufacturing or any sort of machinery that has sensors.
That collects data like the voltage and rotation speed of machine or position or anything like this And they collect it over time. This is a time series data. inherently You can forecast these timeseries but then other examples For example supply chain domain you can try to forecast the demand for certain items, or shipping. Or traffic flow when it comes more industrial like infrastructure settings and admission rates in hospitals.
And also the entire financial industry is revolving around time series with stop prices, option pricing... many, many examples of time series. Okay sounds great yeah and I personally have been involved within the industrial space for a time series specifically. we're going to come to that later as well asking you from my experience More, more words.
just that there's so much information on what a Tyrex II is. Share with us the foundation model. everybody you know talking Foundation models. and maybe if we have been talking Foundation Models since I guess two or three years Then that was they were typically Related to.
maybe let's say words, but you're gonna correct it. But now there related two times here. So what is a foundation model and what specifically? Is the Foundation Model when we talk about time series?
Yes You are very much correcting that. Instead of words, I would say language. But then foundation models definitely originated from the language domain with this whole GPT and HRGPT series is what kind of kicked off this line of research? And then applying this concept of having a foundational model that's really understanding like truly a foundation's of specific domain.
In the case of language models, of course language in general and then when you apply this concept now to time series The model that you want to get out is some Time Series Foundation Model That You can throw any data into Any sort of Time Series that you have collected And you want forecast it And then it will be able to just extract the patterns during the forecasting phase already, and you don't need to again train. It like he would do previously before time series foundation models were taking off in have to train unique model for every timeseries.
You have to be able To forecast as well. now we Have these generic Foundation models that can understand Time Series In general. no with language. And as you said, starting with OpenAI I guess three four years ago they and many other companies scraped the internet kind of used close to all available language As you said that was available To us.
is that then similar? Does foundation imply that at least you have a huge amount of. So in this case, if you say time series foundation it implies Yes, there is very much a good analogy that for language foundation models you really had to scrape this web scale data of billions or trillions of words and tokens sometimes. For time series I would say it's similar setting with one extra caveat that i'm gonna get into in a second.
but yes we do have a large collection of timeseries A couple ten million, ten twenty millions of timeseries that were used during pre-training. But the unique property of Time Series compared to language is that you can generate completely synthetic or artificial time series yourself, where you have full control on what kind of Time-Series do want to generate and how. And this also now very much an active research area in this field On How Do You Generate This Data? What should be put into it?
because then you practically speaking Can Have Infinite Amounts Of Data Which Is Definitely Not In Language. So we do have watch care data extended with these synthetic datasets. Okay, good interesting yeah. Yeah.
Synthetic Data has been around for not sure maybe also a couple of years from at least for me. I know that i've been trying to get my brain around it for long time and I did not always understand that. I think there was a very good case when while We were talking about autonomous driving and somebody was suggesting, yeah you don't want to have the... And then typically maybe it's the corner cases in the corner case of where a child runs across the street.
You'd rather not make that choice. so for me there is use-case. I understood very well. I don't want to go into much detail here, but it's my question always is like if we produce synthetic data.
How can... It's almost turning the world upside down. how can we produce data? If doesn't matter medical financial industry or whatever and still make sure that they are representing our world, because in the end if you're going to be doing inference and that's real-world data right?
So I'm not sure how that works. Yeah i am completely with you on that. first time i heard there is even research on fully synthetically pre trained time series models That are also quite good performance definitely better than what would expect through something which has never seen any real word Time Series. but without going too much in depth of the technicalities.
In general, time series and time-series forecasting you just need to understand what patterns appear in your data at what frequencies they appear do they repeat? Do they maybe change? like first there You have a large peak And then you have like smaller peak because it was the weekend or something Like that and you also have trends. Maybe some things is going up continuously go down continuously Or its changing trend And if you design your synthetic data to have these properties there nicely, all we have to do is just train the model to extract this and being able to recognize them.
Later on when it sees some real-world data at inference time then It will be able to recognise like oh I've actually seen This is a pattern repetition. now i'll just copy it. Basically that's the high level. Okay very good.
Then rf two three more. You know, what we're talking about for me is so full of abundance... of information and words that we use all the time. And again I just want to make sure that the majority of listeners...
Sorry if those have already known it but they are going at the same level. Because then exactly what you do in the next word is zero shot. And a zero-shot means that if I am going to take your algorithm and different from what i would be doing, like ten years ago when I was involved more on this side but nevertheless We would have to train. So we will have to show our algorithm at that time, uh the data that we had and we would have trained The model in front of us with our data.
now zero shot means it. I can just take Tyrex two And i think you said also tyrex one already was a zero short an. immediately It Will Have the capability Of extending of looking into the future where my data is moving too. Is That Correct?
Yes, largely that's the idea. Zero Shot in general is a term used in many places of machine learning now for time series foundation models. what we mean when you say zero shot? Is there specific data that you want to forecast?
most likely The model has not seen it. later on can go bit more in depth perhaps about benchmarking these within benchmarks, it's also a big step where you do zero shot. That means the model has not seen any of the benchmark data that we want to forecast on and in industry as well. It is very important capability because most likely if company wants use this their proprietary data from there.
proprietary sensors Is mostly something that We did not have access during training so we didn't train at all. So really make sure this tydex model can generalize to these unseen settings, making it effectively a zero shot setting. And yes you were also correct that the original Tydex was also operating in the Zero Shot regime? Yeah right okay so is the zero-shot?
Is That The Standard Operating Mode? Can I at all. So as I said, you know i'm gonna have maybe let's not forget in the end how to get access And I'm going to start exactly with that zero shot. So, i'm gonna connect it my sensors giving you the data.
we're going talk about univariate multivariated in a moment and then It will be showing me where my sensor is moving. The question Is Can I? Does it make sense also If I have, let's say one month or whatever a couple of hours depending if it is milliseconds or days. Whatever?
Is It possible doesn't make sense to also train it on my data? Or is it typical that it always works but zero shot? and That's very good question And Of course it depends A lot. But in general i would Say does makes sense to have some sort of fine-tuning because for these foundation models, we want it super general and the same way that a language model is made to be general.
For time series models as well if you already know maybe you don't wanna forecast weather patterns or sales patterns but really apply them to machinery and sensor data then at NXA we can also help with that and it can be further adjusted, so it really focuses on the speed patterns and frequencies that you have in your data. I think i understand that good! I continue with my still base questions but we're gonna get closer to the specifics of the Tyrex II. Still there is one that both Tyrexes and Tyrexx II have in common.
they are both recurrent XLSDM-based foundation models. Now, XLSDM. we didn't talk but that's the. maybe you do want to spend just one or two lines on and then in combination with telling us again what a recurrent model is.
Yes! That's also something that we are sort of very proud here at Linz having this XLSdm architecture fully developed. here. This is Just In Two Lines.
it's a modernization of Seb's original LSTM idea And this is a sequence processing backbone in high level. Now the main difference I would say, and i'm gonna compare these recurrent architectures to transformers because they are sort of the most prevailing architecture for most time series models Is that transformer stores everything In its memory and compares everything To every other timestep basically on our high-level in a recurrent architecture, because we already know that our time series is inherently ordered.
We only want to sort of create an architecture so when it sees something on your time-series or new timestep It just puts into its own memory. and this memory remains constant In size. And all you have to do is manage whether I want to store something Whether i want to forget Something But Because Of This Recurrent Nature Of Processing It Sequentially This means that both the storage cost and processing costs are much lower, more favourable which is then again important in certain deployment scenarios for these models.
Very good! We may be coming back to because I think the question of memory... is a very important one. i think it's at the base already.
just you now told in relation need for running my model if it's sequential or then, uh...if is going to be quadratic. Yeah and not forget that the M in the LSDM original twenty seven-twenty eight years ago Zepp and Dr. Farah Jurgen here in Munich stands from memory as well.
so maybe we come back to that later. now The first major difference. I'm not sure I mean, sequentially if that's the number one you can confirm. maybe there is also a little bit depending on what interest.
You each listen to have. but where Tyrex was dealing with univariate data, TyreX two deals with multivariated data? I believe i know what it is But why don't tell us the difference between Univariates and Multivariat? Yes!
I'm more than happy too And would say this is the biggest change from the first to second generation of Tirex. So, to explain it best I'd say that the original Tirects could look back on their history in the time series you gave them but only for itself. For Tirext too they can not only look back at its own history but also now effectively gather information about other timeseries called covariates or multivariate timeseries meaning Now it just has this whole new axis, this whole dimension of data that can have access to and it can incorporate into its forecasts.
And in the entire forecasting pipeline which is I would say quite a big improvement In many scenarios. with multivariate data we also get into what these scenarios are. This will be quite a difference. Yes, that's the follow up to cope various but then I'm not sure what idea actually.
completely understand it. I thought was as easy as univariate is one variable and multivariates a number of variables like as i've typically been used you know maybe would have whatever hundred signals he will bring them down through mathematical ways may be two times... But its'n't as easier than that! It's not that Tyrex only does looks at variable.
That's not what Univariate means? No, that part you did understand correctly. so univariates mean the only look at one variant as a time and for cost only. based on that history.
very good then I did okay yes Okay yeah Dan Yeah i can only confirm it. You know from my little experience I must say when was involved in several jobs they were typically In an industrial environment be more than one. So I guess maybe many listeners are then also looking forward to be working with Tyrax too. for that reason now you already extended it, which i'm not so certain but past and future covariates.
What does that mean? Exactly, so covariates in general are just additional time series that the domain experts or someone who wants to use it knows they're probably relevant for this target. we want to forecast And past covariates basically only have observed up until as the target. To give a brief example, we want to predict let's say the traffic flow of a junction and we also have real-time data on how much it's raining at the moment near that village or something...
And you know then more likely people are gonna take their cars. they don't wanna get drenched. so then we incorporate this covariate and incorporate it into the forecast that, okay most likely now I will increase my forecast a bit because i know that is going to go...it's raining more at the moment.
On the other hand future covariates are covariates where we also brief example on when this is a perfect application case. So let's say you want to forecast your sales for your company and in retail it's known that before holidays, before Christmas etc. the sales do go up but we know when Christmas happens. This already has set-in time event And We Know It In The Future.
meaning give this information to Tidex as a future known covariate and it's going to know that. okay I know there is this event happening in the future, leading up to this event. The sales will most likely go up. they would drive out so i will also incorporate it into my forecast but they'll increase.
That's the main difference between these future-known cases Today, defining well at least the original choice of what is a potential or a covariate. Is still human activity? Is that correct? I mean are we?
and then the point this but you need to if your first needs be a human too understand What is Christmas And Then You're Gonna Say Okay Christmas and then you say okay December twenty fourth Or Whatever. But Maybe in The Future if the in a specific environment, The model is going to know all these different variables. Maybe then the algorithm would come up with something maybe by not saying by Christmas but say there was some thing I see invariable XYZ which it's relationship two covariates with something else without knowing that its talking about Christmas.
and Yes, that is to some extent yes. You do have to define as a human the potential user of Tyrax on what covariates you want to include. but the way Tyraxx too is built up it can inherently also decide from all those covariates which ones are actually meaningfully contributing and then put different amounts of emphasis. so he does has capability.
Of course, you can't give it thousands or a million different covariates and then just let the figure out. that is a bit unrealistic scenario. But in general if your not sure of this useful or not You can confidently include as well. I do recall exactly, i may have some years ago given this example but it was only after a long time domain experts on the shop floor not understanding why every now and then.
The quality of the layer on a table was having bubbles, whatever. Until they found out it had to do with somebody opening a window in their back because there were smoking I believe and So it had to do with, you could say wetter or a window whatever. Which you then only by adding in new variable what is... It's almost like Is the person smoking on that?
Or is it cold outside and suddenly your gonna see their relationship are doing? You're going find out why Things happen now the tire x two future covariates. they ensure I understand strict target causality. maybe that's a final thing i want you to explain.
We talk many, many times about correlation and causality. when u say that the covariate insured the causality What does that mean specifically? Yeah, sure. That's also a good topic.
so on high level this strict causality especially in the future is important because it of course In time series forecasting. we cannot let any future information leak back into The past as that will just make the target too easy for the model. It's like oh I just repeat what i saw already that we are making sure this future information only enters the model at appropriate time, so there is no leakage of this unwanted future information into target and then causing these unwanted or spurious correlations.
That's a different topic! It's a great book that shows all these previous correlations. Yeah, yeah I mean i could but we're not going to do that! I think i could be talking for rather long time with you on this topic of causality because i have this feeling there is huge potential in you claiming and im sure if u claim thats just take it as it isn't what that means.
uh... But.. We are Not Going To Do That Because ..we do wanna get closer to Tyrax too here.
so now As we talk about I have understood the multivariate. Maybe there were already, today a number of multivariat foundation models in the market and i assume that you are comparing yourself with those maybe without, you know making kind of positive statements or whatever. but there was other models in the market and towards the end. Or from now on we're going to be looking at how is Tyrex too doing?
In relation to these. so maybe he can share some base numbers with us. Yes truly! And yes You are definitely right.
this I would say a couple years ago when first time series foundation model area kicked off everything. Multivariate is definitely a step above and it's more of challenging problem to solve. But now there are others, so probably the most well-known one from Amazon this Kronos II model. This is fully transformer based model that does not have this fully causal structure.
And thats probably one of main differences. And the other main difference is that we haven't really talked about parameter counts for these models explicitly. So I just want to mention it briefly, that there are two models? Yeah Just tell us Tell us what are parameters and What Is The Value Of Maybe The End User The Listener Who's Gonna Decide To Go For One Or The Other Solution.
Also Then Knowing The Amount of The Number of Parameters. Yes Surely so. then Probably. then this other difference lies in the parameter count, and then in connection to this also how efficient can these models be under separate deployment scenarios.
So this parameter count is just the number of individual parameters that model has to learn And all have to stored on their device and load it into a fast memory during inference time. meaning the lower your parameter count is, then more restrictive hardware you can run the model on and also quicker you get your forecast out which if you deploy it in a scenario where you have data coming every second this very important. So for Tirex, the multivariate model of TireX-II is eighty two million parameters And Chronos II has one hundred twenty millions.
so there's I would say an advantage of Tyrex too. But what's also something unique is if somebody decides that they do not need the multivariate capabilities of Tyres, you can turn off those sets of parameters and then this just gets reduced to thirty eight million which has been a lot smaller on enables us A lot more efficient. Okay now it was something I learned. maybe i never realized Maybe when you train these models weeks or months, right?
And depending on what kind of CPU you have available. You gave us the example of the hundred twenty million Kronos to eighty two or thirty eight million. for Tyrex. to multivariate a single variant is say so that model in the end is the contains all these parameters and you load them where you load him in what memory in CPU memory are a time of inference I mean?
Yeah, that kind of depends on the application scenario. Of course in the case of Tyrax we see many applications were on edge devices which have very limited compute capability. so then you would try to load it into RAM or just some fast memory access. If your GPU is available and you want to do inference on this then these parameters will live.
Okay, typically. Right now what? What is the number that we need to think of them? I mean so for each parameter you take whatever points something megabytes.
So let's say for an average of a hundred million parameters You need what kind? if? are we talking gigabytes or were talking terabytes? I would say these are relatively still small scale and then with some clever engineering tricks you can shrink this down to, i would say under a hundred megabytes for something like Tirex at least where...
I'm bringing up TireX of course because we know a bit more on how it behaves under the circumstances. but i'm sure these tricks will also work on generally any model. So it's not that huge compared to something like language models where you have hundreds of gigabytes or terabytes of data just in the model parameters. Very good, so when maybe a little bit later going be sharing some specific numbers we were now comparing kind of features are capabilities I think and understand.
if for see yourself compared to understand that both multivariate, they boast through the past co-convariant also. The future covariates. what about streaming? Yes!
That's a very good point. so streaming is in high level. if this is something I think easiest to visualize with an example from industry let say we are collecting data from our sensors and this sensor is measuring rotation speed of machine. but this rotation speed changes every second.
We get new data, uh, every second and we want to then continuously forecast what a weekend do with Tilex too? And these recurrent architecture in this fully causal structure that we built is we process the history that we have so far and redo our forecasting. And then whenever we have now a new observation come in like Now we know again one second of data or five seconds off data. We can just do a single quick step of, okay now we put this new information into our fixed memory enabled by this recurrent architecture and you just do another forecast.
Meanwhile on something like chronos two or any other transformer based model in general that is not fully causal he would have to recompute A lot more off the past than sometimes your whole past again Just because you added a tiny bit more information. This Is also Something That's heavily can influence the inference speed and how quick you can get your forecast in these new scenarios. Because that would again be quadratic, as you said before rather than sequential? Exactly it's quadratic!
And also just generally we have to then re-compute everything again and again. We don't have to recompute things That already did Just add things. Then uh...we tested this to quite a long horizon This streaming capability.
It remains stable To quiet large extrapolation length. Now, is there any limitation? you gave the example of getting new value for single multivariate variables on a second level? Many times within the industry we have millisecond deterministic values coming in.
Is their limitation from the Tyrix II algorithmic perspective or limitation that sits more in the architecture of, you know CPU GPU access to memory etc. You mean like for when it comes to going below second. so if every now and then you do the calculation. So, you have the capability of this streaming?
Can that streaming go down also on a millisecond level? so I get one thousand values per second?" Of course yes okay no i understand...of course there will be hardware limitations as well.
if you have data avenue millisecond i would say at that point just purely loading it into the memory already take longer than a millisecond. So you can do this update, but you can go I think quite a lot down with these Okay? And one more. i'm going to be asking you as I said before that the importance of memory we're really talking.
so what about? I think Tyrex two understand users has a constant memory. What does it mean exactly? and uh...
what is maybe then the advantage in relation? i can certainly get into that. so this constant memory is specific to these recurrent architectures compared to transformers and of course then tyrex in this case representing the three current architectures. chronos too would represent transformer architectures.
So what transformers do, every time you get a new data point it does some computation on it but And whenever some future value comes in, it will then compare to that data point again. But as you add more and more data points You have more things to compare which then grows quadratically quite rapidly Whereas this constant memory setup basically has a fixed size of memory As a vector or matrix. mathematically The key idea is only having the fixed amount storage and the model just learns to decide, okay what do I need to store?
What is it that i don't need to Store at this point. But This also means That we are fixing The state size of the memory in the beginning And It cannot grow further than that. even if you were To compute for a million time steps You will have the exact same megabytes Of memory occupied on your system If you compare Just two ten times steps whereas with a transformer very, very quickly into something that's not really manageable unless you have a high grade hardware. Okay now share some.
I don't know how you want to do it? Some numbers? Well number one i think was already the case with Tyrax is exactly this relationship or difference between the sequential transformer quadratic so its more energy and CPU. We know that, and that is the same.
That relationship with Tyrex too in relationships those solutions based on transformer like understanding Kronos to do. you have any other specific? Is there only any others specific numbers maybe just supporting what it is? we talked about so many I assume I don't know them, but as many different benchmarks.
As always is there one or two benchmarks? You say those are the ones that you and other providers of the time series foundation models are looking at And our listeners are looking At for which reason they're looking A or B. Where is Tyrex II in relationship maybe to competitive models? Yes, I can definitely adjust that as well.
So when it comes to these benchmarks for Time Series Foundation models the overarching goal is test on a wide range of different domains and settings. so we really see which time series models are truly generalizing. In this regard i would say there's two that you could bring up here. one is GiftEval Which has been around now and where Tirex-one also rang at the time of its release, to this day actually quite much on top off the leaderboard.
Here we do not fully test for these future known covariates cases or did... This doesn't test so heavily on the on the covariate side of things but For example it tests for a longer context. Or he tests very generic ability to forecast in different domains. And then this other one is Fevbench which is then much more geared towards having these covariates, this past only covariates.
These future known covariate cases, these multivariate cases and I can say that in both of these Tyrax too it's ranking very much on top with the best state-of-the art models, with these timescale foundation models that are often quite a bit larger in parameter count. I'm talking hundreds of millions or now. we're also seeing billions of parameters and we are very much competing with those models at a much smaller scale. as i said eighty million parameters what you were operating.
okay. so would you suggest listeners decision makers? That our interest is evaluating that are not yet inside of these benchmarks. So should they be looking specifically?
You know, the majority of the listeners will be in an industrial setting and if it's an industrial sitting on we have typically I would think multi-variate. but who knows maybe some of you do look at or interested in looking still is then one or the other benchmark The more specific. if you are going to be talking to people in the financial market, maybe seasonality retail another benchmark would be more relevant for them. So I would say in general because one of the largest improvement of Tilex too is this multivariate and covariant support And Fevbench very much designed To do.
looking at that will give a better impression on how well we're doing when it comes to these covariate cases, this future known covariate case especially that only Fairbench tests for. So I would say is a good place to start in general and of course trying it out on your own data is also usually very nice way to get the good feel about how actually does behave. Okay! That's a good point.
now How does that work? Let us say more importantly one of our listeners says yeah thats exactly what i'd like you do. Can I test? Can I get an evaluation version?
can i Get my data to you or do You provide me with a Tyrex two and I can test it myself. And then, I compare It too. whatever other may be possibly interesting for Me as well. how would that work?
Yeah, so much like the original Titex model where it was available fully openly on Huggingface. TiteX-II is also available with very similar licensing system and a very similar term but everyone... The main goal for us is that everybody can try out and see what they could do and of course for any sort. If you then like it or if you think that maybe we can make somehow even better for your use case, then NXCI is definitely here to help And they should get in contact with us as well.
As the base model is very much available Very good rounding it up. Maybe you know, what have we just heard from You in the last? oh? We already talked.
Peter always talks for about an hour with people that I have in my In my podcast bringing it together. What is in one two lines? What is Tyrex to? more than other then what tyrex has been until today?
So, Tyrex-II is just this next step of the previous TyreX model that can now incorporate other covariates in its forecast. It can do multivariate forecasting natively retaining very much similar capabilities with regards to this constant memory and scaling. And there are also these multivariat improvements which make it even more competitive in this benchmark and the landscape. Sounds great, before we close Levent tell us you mentioned that you are based yourself I believe?
That's what you said in Linz. Tell me a little bit about maybe also assuming as a company looking for talent how does your work or day working on product like Tyrex too. How does it look? What is the kind of things that you work on and what may be type people talent where as a company, do not yet have enough off?
Yes! That's great question. so I am part of research team also PhD students my day much a blend of what a PhD student would do, reading papers catching up on new research and then also gathering ideas. But then on the other hand NXAI's domains like this time series domain especially is then heavily influencing my research than I'm trying to work on.
things that are done can be made into NXAI business in general. That's what we do under on the research side. And then of course we have very, very good engineers that can translate these models that will develop into actual customer projects. So I would say if you are interested in that and working with this time series model or even more so getting them to relive the employment scenario sometimes challenging scenarios than they should definitely get in contact us.
Okay, and typically the people around you yourself are people that have come from what? From mathematics. You know starting with SEPP I know we started here in Mathematics... I happened to be at Munich Technical University again last week for other reasons but i thought of SEPPP at that time.
so is it typically mathematics as these days? maybe you tell us One, two words about how it works in lenses that typically damn the what is at a direction of I don't know artificial intelligence machine learning or watch. What typically would you be? looking for people, bringing with them if they would apply or possibly working.
So the backgrounds of two people are definitely from many different places. but If you studied something like computer science or data science Or worked in the past machine learning projects With maybe actually hardware level optimizations even or like a hardware engineer, that would definitely be something that fully fits in. But then yes as you said also we have people from mathematics and different backgrounds. so it's not just fully restricted to computer science Or this machine learning and AI area of studies and expertise.
Very good! And not to forget the mention being close around one of top researchers in this field, Zapp Hochreiterm which I'm sure is wonderful for working with. Levente thank you very much for your time. Thank You very much, Cherng.
with us. The details of Tyrex II. Thank you very Much and Good Luck. And I would say those listens of you that are interested, yeah.
You know where to go. Enix AI. maybe you can contact the vendor directly or somebody else at Enix. Yes, but I'm happy too.
Thank you very much. Thanks for having me. Bye bye.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.