AI Speech Technologies π€ for Business π Productivity π and growth!
Dave Erickson 0:03
Your receptionist just quit. Your customer service guy misunderstood the voicemail and shipped the wrong parts. It's enough to make you want to replace them with a robot that speaks. On this ScreamingBox podcast, we are going to discuss AI speech technologies. Please like our podcast and subscribe to our channel to get notified when the next podcast is released.
Dave Erickson 0:44
You're worried that using AI speech to text for your business will make your communication feel robotic and full of errors. No, I ordered ten cats, not ten bats. The reality is AI is quietly transforming human facing and communication oriented parts of many businesses. Welcome to the ScreamingBox Technology and Business Rundown Podcast. In this episode I, Dave Erickson, and my human speech oriented co-host Botond Seres will explore how AI speech to text and text-to-speech technologies are redefining what it means to interact with a business with Gene Sigolov, co-founder of Upfirst.ai. Gene is a serial entrepreneur, attorney, and technologist. Before Up First, he founded Simple Texting, a business SMS platform that became one of the most widely used messaging tools in the US, serving 17,000 customers. He bootstrapped simple texting from zero to 20 million ARR before selling it for nine figures.His multidisciplinary background, a BS in computer science and a JD in law, gives him a rare ability to blend technological insight with real world business and legal experience. As co-founder of Upfirst.ai, he has worked to build an AI powered receptionist platform built to replace the traditional virtual front desk staff with a smarter, more capable, and a dramatically more cost effective solution. Jean, welcome to the podcast.
Gene Sigalov 02:10
Thank you. Glad to be here. Thank you very much for that warm introduction.
Dave Erickson 02:14
To begin with, what came first, being a lawyer or a technologist?
Gene Sigalov 02:19
Truly being a technologist. I've loved my, I've loved computers my entire life. My first computer was an IBM PS2. It had, you know, half a megabyte of memory, 20 megabyte hard drive, and I loved that darn thing. Right? It only ran MS DOS, was not powerful enough to run Windows, no no version of Windows. And yeah, that was the beginning for me. I was probably 11 years old and my father got that computer by way of barter. Someone had owed him money for some sort of work and instead of money he got a computer, which means I got a computer.
Dave Erickson 03:00
When you started simple texting, was that kind of before you were interested in law or you had already gotten your law degree and were trying to practice law or what drove you into that?
Gene Sigalov 03:11
Yeah, yeah, it's it's it's it's a it's a great question. I'd already gotten my law degree, you know. I went to law school without the, necessarily, without the intention of ever practicing law. I'd actually intended to go to business school and I talked to my dad about going to business school. He said, Look, son, you know, if you look at the world, a law degree is really a very utilitarian choice. Many people in business and politics have law degrees. It doesn't preclude you from doing anything, but it gives you the ability to practice law should you ever need to in your life. While a business degree doesn't really give you the ability to practice anything, right? It may be a useful discipline to study, but it doesn't give you any sort of privilege, whereas law does. And so I went to law school.
Dave Erickson 03:57
Law is the foundation of a lot of business mechanics. And if you understand how the law works, a lot of business mechanics have to follow the law. So that's very helpful in business. I have done a lot in, in writing contracts and licensing agreements and all that, and I found that studying law was very helpful in doing that, even though I didn't get a law degree. I've studied a lot of law over the years, just 'cause that's how you need to to write contracts and licensing agreements.
Gene Sigalov 04:25
Surely. I also, you know, I compare it a lot to like a, a really good social science degree because you really understand the mechanics of human agreement, right? Like what's the default, what are people's expectations, and you really know how to circumnavigate life without falling into any pitholes, legal pitholes, right? So it really gives you an understanding for how to structure deals, how to deal with people, just all the things.
Botond Seres 04:53
And also on the utilitarian side, doing a degree that allows you to practice something, anything really, is a boom in life in my opinion. I found that out the hard way. I did a software engineering degree and as I found out later in life, that doesn't give you any privileges because apparently anyone can do software engineering.
Dave Erickson 05:17
When you started Simple Texting, you obviously got a background in texting technologies, messaging technologies.after you sold Simple Texting, did you naturally drift into text to speech, speech to text technologies or did you go somewhere else and later discovered those technologies?
Gene Sigalov 05:39
No, I had worked for the Incquirer for a while, for probably a couple of years, but I had this idea that AI was sufficiently good to do a receptionist that people would open their wallets and pay for. And I kind of tried to Swat the idea away for about a year, but it just wouldn't go away. And see, I'm also a consumer of a receptionist service, right? I have a law practice and I was using a human receptionist service, which is pretty expensive. It was like 300 bucks a month for 30 phone calls. And each additional phone call was $10. And so I've also used a receptionist service in the past while building Simple Texting. We had used a receptionist service. We had like a little bit of a liminal period between when our support staff in the US and the Philippines would be offline. And to fill that liminal space, we used the receptionist service, which is also again expensive. And I thought that, and and by the way, these services, the virtual reception services, they don't know a tremendous amount about your business, right? They're following the script, they pick up the phone, you're picking up the phone for so-and-so. Your job is to collect their first name, last name, reason for their call, who they'd like to speak to, and then to summarize that and shoot it off to someone. And those experiences aren't stupendous. And as we're now learning by, by speaking with clients, you know, you would think that people are gravitating towards a service because they couldn't have afforded a virtual receptionist service before and now they can afford the service for like 25, 30 bucks, right? So now the it people, people who would would would have been without a service are now getting some service. That's not entirely the case. People who are tired of their receptionist service are gravitating towards AI as well because the AI has much more context about the business than a person does who doesn't know anything about your business. He's just opening up, you know his screen that says you are answering the phone for so and so your jobs do so and so and so he only has a couple of sentences of context versus the eye, which can have, you know, huge voluminous context.
Botond Seres 07:35
Gene, I have to be honest. I didn't even know that there were virtual receptionist services. I was under the impression that if you get a receptionist, you get a dedicated one who works for your business exclusively.
Gene Sigalov 07:47
No, it's th I mean it's, it's very expensive, right? To have like a full time person, especially in the US, right? Like so the the thing about receptionist services in the U, in the US is that number one, they're expensive and these people don't want to be a receptionist for their entire life most of the time, right? So they might be a student, they may be doing something, this may be a transitional role for them. And once they get their stuff together, they move on to something else. And so low wage role that really doesn't gratify anyone, not the role holder, not the people who they're servicing. So I I I I think that this industry is going to get eaten up by AI, just a matter of time, right? That's kind of like a secular shift. is it gonna take five years, gonna take ten years? I can't say. there's certainly going to be some holdouts who are like, you know, hey, we're providing you a human touch, a human service, you know, there's no AI. You know, surely people will try to distinguish them, themselves in this way, but the economics are so strong. and not just the economics, also the level of service that you can provide with AI, that I think that the vast majority of this kind of role is going to get eaten up.
Botond Seres 08:57
Absolutely, I agree. Maybe I would argue that if you have a dedicated receptionist, the technology is not quite there, but for a virtual one, I do think I have to agree with you.
Gene Sigalov 09:10
I would agree with you a hundred percent. And I would even go further that, you know, for some high ticket items, you do want that human touch, right? You know, I would be upset if I was shopping for a Ferrari and a virtual receptionist that was an AI would pick up the phone. I would just, you know, I would probably think less of the brand. even, even, even less than I think now after their electrical car release.
Dave Erickson 09:35
Yeah. You mean the Apple Ferrari?
Botond Seres 09:36
you
Gene Sigalov 09:39
The Apple Ferrari, the John Ivey Ferrari, John Ive, yeah.
Dave Erickson 09:43
Yeah.
Botond Seres 09:46
I admit I forgot about that. All together.
Dave Erickson 09:49
I did a lot of business in Asia and Europe. I traveled to Asia quite often and was involved in e-sports and gaming. And the topic of speech to text and text to speech was actually fairly active. A lot of it because of translation, because of data entry. Doing data entry in different languages, particularly ones that are character based, like Chinese or Japanese, is kind of difficult. And so they were having a lot of conversations about, you know, speech to text, text to speech, dragon came out, everybody was kind of excited about that. They use it for some rough translation and stuff like that.
Gene Sigalov 10:33
I I'm, old I I'm old enough to remember kind of the genesis of Dragon Dictate, right? This this product was like version twenty or or god knows what versions probably been out since like the early aughts, maybe even l late nineties. It was quite interesting, right? Never quite good enough, but always very interesting.
Dave Erickson 10:50
Yeah, and, and that was kind of the barrier. And I I I think that, you know, even up to a couple of years ago, four or five years ago, the concept of text to speech, speech to text, it still was very rough. You know, they were happy to get 70% accuracy. And so there was a lot of those barriers. I think AI has helped that significantly. And also the push by the automakers to bring
audio text and and other commands into the car, particularly Tesla, who's done a, a pretty good job of that, accelerated speech to text technology and and that. How did you, as you were looking at this receptionist or upstart AI to begin with, how did you go about, like looking at what technologies to do for speech to text? How did you evaluate those and where do you think it is as of today compared to a couple of years ago?
Gene Sigalov 11:49
So much to say here. so with regard to a couple of years ago, I would even go further to say that people have gotten kinda burnt out on the general technological area by dint of some solutions not being quite there yet, solutions like Dragon Dictate, but also intent based systems like for example Dialogflow, which many of the large companies like Apple would use to answer the phone and they would do intent-based routing. So they would ask you a question, right? Maybe give you four choices, determine your intent, and then route you further down the phone tree, right? And it kind of felt like AI, but it never felt quite good. And I think that many people, as these new Alum-based systems came on the market, kind of confused and are maybe still confusing them for this old intent based system, right? Which kind of feels like AI but is not really, versus this AI, which is really AI, right? So I think that's point one in terms of how we compared and contrasted the offerings that the market had at, at the time when we were getting started. And you know, this is an ongoing dynamic process, right? We re-evaluate the state of the art at any given time and we make a determination whether or not the state of the art has evolved to where we need to use a different vendor for part of our stack or we should stay put. And it's kind of complicated because you, you want to be on a stack that has a good future, that has a lot of investment behind it, right? But you're also, there's always the new shiny object, which has maybe just a few millisecond faster speed, right? You know, and and it's it's it's also it's always a balance, right? There's obviously a transactional cost to switching stacks, switching technologies, right? Instead of furthering your, deepening your platform, you're now spending time switching technologies. So I mean th, those are all the facts we that, that informed our decision making. You know, what's the state of the art now? What's the future of the platforms that we choose? So those are some of the things. And you know, so can
Dave Erickson 13:56
I I mean I can tell when I get on the phone with something, I can tell intent immediately and they are the worst. I I I mean a lot of banks use them, my banks use them.
Gene Sigalov 14:07
Yes, they're terrible. They're terrible. Human please, human please, human please, human please, right?
Botond Seres 14:13
But it doesn't work. It's the systems I interacted with, it no longer works. It just keeps repeating the same statement over and over again.
Gene Sigalov 14:22
Yeah. And and so but but the y you know, to to to your point, I I do think that there was like at least at the outset when I got started a couple of years ago, year and a half ago, there was a lot of confusion between these intent based systems and what we're doing now, right? Because people people who aren't like online all the time and aren't really into technology, they don't really know, right? and so, you know, that's still remains, right? There's still some confusion.
Dave Erickson 14:45
The. the keys to AI is, is controlling the AI in a sense. So if you start a conversation with AI, it can go anywhere. And AI likes filling in details and likes adding details and likes, you know, it it in the background, whether you believe it or not, many AIs have programming that says please stimulate conversation so it uses up more tokens. so they can be very chatty and, and go in a lot of directions.
Dave Erickson 15:13
But with a receptionist, you kinda wanna keep the AI focused a little bit. So how did you approach that, that part of AI?
Gene Sigalov 15:24
Prompt engineering, I mean, that's that's really all there is to it. because you know, I serviced my use case first, which was the attorney use case. It needed to be really buttoned up, right? So no volunteering of information, no con can conf, confabulations, no confabulations. so I started it with my own use case first and we buttoned it up tight, and we in fact had to kind of release it a little bit, and you know, if you have more intelligence also, there's less ca, confabulation. So right now, if you're using open AI, for example, which we are, you can choose the level of intelligence. It slows things down if you give it more intelligence, right? Higher intelligence, but of course it improves the quality and it's always that balance.
Botond Seres 16:11
Right. Gene, would it be okay if I got a bit technical here? Yeah, course. So you, as someone who has worked with AI systems and prompt engineering quite a bit, I assume, how long would it take you to jailbreak one of these systems? Because I do see this trend on social media. Everybody is jailbreaking AIs for fun and profits. But I haven't really heard an expert opinion on this, how have the cartridge evolved?
Gene Sigalov 16:42
I don't think I can personally do it at this point in time. Maybe I could do it a couple of years ago, but I think that the Frontier Lab companies have gotten very savvy and the people who are jailbreaking these systems are very sophisticated actors. I don't think it's, it's a part time hobby any longer for someone like me. So I don't think that you know, you, you flatter me. First of all, thank you for saying that. But I I don't think that I can I can jailbreak it. And I try occasionally, like, you know, t take I don't know, Claude, right? Fable, right? it's so buttoned down that even if you ask Fable some medical questions, right? it'll downgrade to Opus because it doesn't want to they, they haven't figured out the controls for preventing you from using it for like frontier type type medical research versus answering your own medical questions with your own medical record, right? So they've erred on the side of their own personal fun and profit versus you know usability. So I, I do think, think these systems are, are very buttoned up now, especially the more sophisticated ones.
Botond Seres 17:49
Right, mean, we did get a report that about 90 plus percent of the code for these systems is just Gardele's now.
Gene Sigalov 17:56
Totally believable. you know, we also got reports that OpenAI is hacking systems and, and breaking out of the sandboxes, right? So, you know, it feels, feels you know, and it's hard, of course, to tell what, whether this is like, sensationalistic stuff for them to kind of bolster their brand and to, you know, how cutting-edge they are. But I I think it has a lot of truthiness to it. I think it's more likely than not true. These systems are very sophisticated. They're at this point much smarter than I am, and you know, think you know, I I think their memory is worse than my memory, right? I think that they, they still have trouble with memory. but in, in terms of solving a discrete problem, which is memory bound, which doesn't have like, an extravagant corpus of context to solve the problem, it's just much better. It's much better at everything than me, right? It's better at math, it's better at biology, it's just better.
Dave Erickson 18:48
The memory issue, a lot of that can be solved by making sure you have lots of contacts in their context files because and, and have them read that at a lower rate so you aren't paying so much token wise. But that'll help with the memory problem. But a, a lot of people don't do that, so
Gene Sigalov 19:08
I agree, right? There's also a concept called progressive disclosure, right? Where it knows what to progressively invoke and read into context and when. even with all those kind of useful things, it still it still memory holds things, I find. And I'm I'm a pretty sophisticated user. Like I'll do like a task, like I'll have sixty, eighty agents run, and it's not perfect. It's very good. It's, it's incredible right. So superlit superlatively incredible, right? I c, I can't, I can't exaggerate how incredible it is, but when you're using it all the time, obviously you see kind of the contours of what's possible, what's almost possible, and what's not quite possible yet it's good but not perfect. You know, God, God willing it'll get perfect, right? you know, I'm an optimist, right? You know, the idea of curing diseases with this technology is, you know, reducing human suffering is one of the most tantalizing things.
Dave Erickson 20:04
Yeah, I did some of the medical trade shows recently and I've seen some pretty incredible stuff with AI doing diagnostics on different scans and other things. And predictive outcomes are also something it, it seems to be doing quite well at. But they have to be very careful because of the hallucinations and other things, to check everything. But I think it has some real potential for that.
Gene Sigalov 20:32
Yeah, I I'm I mean, i i it it has some potential at basic, I think, scientific discoveries, right? it's already solving some math problems. I think it's doing some lab work, right, where it's like manipulating lab equipment to perform biological experiments. I I do think that OpenAI is doing some of that work and they published it. They have some web pages dedicated to the work they're doing.
Dave Erickson 20:56
Iβd like to go back to a topic that you mentioned, and that was speed versus quality. And with text-to-speech and speech to text, speed is a definite component. And maybe you can talk a little bit about in your efforts to find that balance, what have you found out about speed versus quality? And what are kind of, some of the profiles of those who might want to go more for speed or more for quality.
Gene Sigalov 21:27
Yeah. I don't so I'll tackle the first part first. so speed versus quality, right? So and I I'll kind of maybe illustrate based on how we've implemented platform. So we use OpenAI and we open AI uses Whisper to do speech to text, right? That's their technology. And so Whisper, if you read the transcript of the real time transcript that Whisper produces, it looks kind of like GoblyGup, right? It's good enough to have the conversation for the AI to understand contextually what the conversation's about, but when a human reads a transcript, it's no good. Okay. So we use the open AI transcript, which is real time, but we don't show it to the user. We have like post processing using Deepgram, right? To provide a much higher quality transcript, which is then reviewable by the user post phone call. So that's kind of like the speed versus quality, right? You need high speed to do the thing, right? To have the conversation, but then the presentation layer for the user, there's post there's post-processing work to show a higher quality transcript. And OpenAI is also taking that sort of approach, right? They have a new duplex speech-to-text and text-to-speech system, which they don't have an API for yet, but they they've released it on their app. And that uses also this kind of approach. They have like a real time Whisper and they have like a post processing and they're going to release the API soon. And you know, we'll give it a try. We'll see if it's as good or or better and it it should be it should be a better technology stack than what we're using now, because right now we're using their real-time API, which is quite good. It uses this thing called semantic turn detection, right? The way that determines that it's the AI's turn to speak is that based on the content of the conversation, it makes a determination that like you've finished the thought, it's it's turn to speak now. but now like it's a full duplex and so it's supposed to be always incrementally making that determination. It's supposed to be a little better. We'll we'll see the the judges out. but to to to to get back to to your your question, so that's kind of like the speed versus quality issue, just like on the speech to text, but then there's also like the actual, you know, the intelligence of the speaking, right? And you know, we've erred on the side of lower intelligence, right? That's the default. and having a more fluid conversation. because when a person comes in for a trial, you want to give them a good experience. And frankly, the modern models, right? The ChatGPT five plus is so good that at basic conversations you don't need a high level of intelligence, right? It's not looking at a brain MRI. It just needs to be like appointment stuff, you know, what, what time's your office open until? You know, do you, do you, do you take care of lions or do you only handle dogs and cats, right? You know, vet veterinary medicine, believe it or not, they, they do get questions about exotic animals a lot. So that's just a funny example.
Botond Seres 24:29
Now that you mentioned it, Gene, actually the only gripe I have right now about day-to-day just for fun conversations with AI is the part where it tries to figure out when it's turn to speak. And when I finish my thoughts, I feel that it's not quite there yet, but compared to like 5, 10 years ago when...my main gripe would have been that it mistyped literally everything I said. It's a massive generational improvement. To tie into that topic of improvement, I was wondering if there are some areas of improvement that you would like to see these companies work on to make your product better.
Gene Sigalov 25:19
I mean, that's one of the big areas of improvement is determining when it's the AI's turn to speak, right? You could improve the determination that it's able to make by increasing latency, right? But if you want to keep latency low, right, that is something that it may have trouble with. So I think that's a design decision that the consumer Chat GPT app may have made is to err on the side of this determining that it's the AI's turn to speak at the risk of talking over you, but they could improve it by increasing the latency. So I think a part of it is a design decision, but hopefully a design decision that they won't have to make in the future once the tech gets better. So I think latency improvements and all areas of the stack would be better, right? The telephony stack, the transcription stack, just everywhere, right? Fluid conversations are very fast, people
take milliseconds to make a decision on what to say or when to speak. So I think just all those little improv, improvements which will lead to a feeling of speaking to a person. So doubles in the details.
Botond Seres 26:31
Maybe try to prompt engineer on latency, like tell the AI to wait more or wait less.
Gene Sigalov 26:39
Yes, w this is kind of early experimentation, but it it it it doesn't help. Not in my experience. Yeah. The the level of intelligence helps, but not not anything else.
Botond Seres 26:51
And is that the level of intelligence, is it something that you get to configure on a dashboard or is that something you prompt engineer?
Gene Sigalov 27:00
no, it's a configuration option. It's a configuration option for the OpenAI real time API. I think it's two point zero plus API, has a flag for thinking and default is low.
Botond Seres 27:15
Right, so if they provided an option for latency, that would be a big improvement, right?
Gene Sigalov 27:21
Yeah, but you can't I I I I don't think, would you like it to be slower, would you like it to be faster? No no no no no one would ever choose latency worse, right? So the the toggle's not actually latency, the the the the toggle's intelligence, right? But that that's kind of like the proxy, I guess, for latency. So
Dave Erickson 27:39
Maybe we can just map out a little bit for those who are looking at speech to text, text to speech, in, in applying it to some business model. The components of it are you're you're listening, so you're converting speech to text. Then once that conversion has happened, you need to analyze it to determine what they're actually asking.
Gene Sigalov 27:57
Yes.
Dave Erickson 28:06
And then you need to come up with an a, a response and then you need to convert that response which is technically in text to speech and then output it. Is that kind of the basic flow?
Gene Sigalov 28:20
That's kind of the basics. Yeah. you know, if you're using OpenAI, for example, you can use one of their voices, but they don't really have a rich library of voices. So we, for example, use Eleven Labs, which does have a, a rich library of voices. You can do voice cloning and you know, if you want like, a southern voice that sounds like it's a native Alabama dialect or accent, you can get that and that's your clientele and that's how you want to present yourself. And so yeah. But, yes, that's essentially it.
Dave Erickson 28:51
What if you wanted a, as more devices include video as well as audio, a talking head as part of the reception package? Is that something that you've explored or looked at or what it would take to do that?
Gene Sigalov 29:07
No, I've, I've thought about it. I've not explored it, right? You know, people don't have video phones, right? So people are still using phones as phones, right? So you, you're probably thinking about an avatar, right, that is for a video call for a Zoom presentation or some such. And that's not the world in which we operate. So I've not thought about it. But I have a pretty good friend who has thought about it because he's working in the recruiting space, right? And he's using AI to help recruiters or companies that need to recruit employees, qualify employees, right? So they need to interview them. And you know, should the interview, the interview is by Zoom usually or a Zoom like equivalent. And so should there be a talking head avatar or is voice sufficient to ask him some screening questions, right? so that's something that, that he's thought about and they've not decided to, to go any way in that direction yet. So I don't think that it's very practical, necessary or necessarily or indeed necessary for a use case which can support an avatar presentation. And also I think that anything that you'd get at this point would be creepy, right? Because it's kind of like you get that uncanny valley effect, right? Like when where it's like a robot and looks really human, but you still know it's a robot, and you get that creepy effect, right? So it's called the uncanny valley. And I think we'll be in that valley for quite some time and you know, at at some point it'll be indistinguishable, presumably.
Dave Erickson 30:38
Yeah, I'm seeing definite improvements in it and we've worked with we had a guest who was with V Dio and they do that and I've actually done it myself of my of my own self cloned and although it isn't perfect I can I can see how it could be used but at there's limits, right? And a lot of those limits are time and other things 'cause it's very hard to do this. The only reason I ask is is not because I think the current generation has a lot of use for a talking head receptionist, but I also have a 13-year-old daughter, and all of her and her communications with her friends are all on FaceTime, on Zoom. She doesn't want to do just voice. She, she wants video and everything when she's communicating with somebody. And she even said to me, you know, I asked her, You should call the school and find out when this thing is happening, and she's like, I don't want to call, I want to talk to somebody. I'd re you know, they don't do FaceTime at school. And so I, I think it's more for the current kids who are growing up and entering the workforce that are, they seem to be very focused on that.
Gene Sigalov 31:52
Yeah. it it it's interesting. Yeah, I have I have daughters as well. They're a little bit younger. I have two daughters, nine and eleven, and they also mostly use FaceTime, but I think it's only because I've not given phones on the regular. So they only have iPads, and so it's their only way to communicate with their friends 'cause they don't actually have phone numbers yet and I don't know if I mean, does your daughter have a phone number? Is she is she permitted to use a cell phone?
Dave Erickson 32:21
Well, she has an Apple Watch so she uses la her iPad or laptop to do that. But that's the, that's kind of the thing. She says a lot of her friends, they, they use it on their phone. So they're, they're, instead of calling somebody just FaceTime from their phone.
Gene Sigalov 32:33
Yeah, Yeah, Yeah. Well I'll I'll I'll know when I cross that bridge, right? I'm trying to keep I'm I'm trying to keep the phones away from from my girls for as long as possible. It's only like they they get phones when you go on vacation or something in case they get lost. They they have Apple watches. I don't I don't consider that a phone. That's kind of like a safety device, right? That they get lost so they can call their parents and such. But you know, they're having a a device free summer now in summer camp, which is a very beautiful thing. We have to if if we write them letters, they get printed out letters of the letters we write and they handwrite us letters. Man, penmanship is so bad now. I I I don't, I don't think it I don't think kids learn penmanship. I'm amazed my kids actually know how to write it all, seeing what they write to me from camp. but yeah, so I I I I don't I don't know. I mean, certainly the technology is getting much better, but it's computationally expensive, right? To do like an avatar of yourself. It's computationally expensive. You can generate short clips. to do like real time conversations, you need probably orders of magnitude improvement in, in compute, I have no idea, to to be quite honest. I would you know, my, my point of reference is using something like Seed Dance, right? Like C Seed Dance version two point five I think is just released, which is probably cutting edge video generation, right? and I don't, I don't think it's cheap. I don't think you can, you can pencil, pencil it for regular conversations just yet.
Botond Seres 34:11
And just circle back a second to the whole FaceTime discussion.I mean, I have a conspiracy theory on this. And I think the reason video calling is still not the default is that there is no network as ubiquitous as just the mobile cell phone network that we have for video calls. I mean, we do have the mobile internet, sure. But if you want to call someone with video, you don't have to do FaceTime, which is exclusive to Apple, or you have to do Messenger, which is exclusive to Facebook. And it may or may not deliver your notification, or you can do Snapchat, or you can do whatever app, but there is no surefire way to say that, I want to call this phone number and start the video call. Still, in 2026.
Gene Sigalov 35:02
This is true. Voice is also optimized for real time communication that's error free, right? So it's you know, decades of experience in providing flawless conversations to people. And I don't know that there is anyone that can provide flawless conversations to people across the globe, right? And a single protocol, right? Like you mentioned. but
Dave Erickson 35:25
Uh, Weβre not having a flawless conversation now, so
Gene Sigalov 35:28
Yeah yeah yeah yeah. I mean put, put point in case, right? you, you and I, right, based in the US are having a pretty flawless conversation, right? But those underwater cables going to Hungary, not so much perhaps.
Botond Seres 35:41
Oh yeah, there's tons of latency. That's why I try to not speak over you guys. I have to wait a lot to make sure you're finished talking.
Gene Sigalov 35:48
It's okay. E E Elon's gonna save us with Worldwide Space Link.
Botond Seres 35:54
Oh yeah, that won't have burst latency at all. If the signal has to go to space and then back.
Gene Sigalov 36:01
I mean the latency is actually pretty good, right? Like there's, there's a lot of airlines now that use it. Also over here in South Florida we have the Bright Line Train which uses it and latency is like it's less than a hundred milliseconds. I mean it's good enough for, for real time conversations, It's like maybe a hundred, one hundred twenty second latency 'cause I'm I'm I'm I'm a internet quality hound. Whenever I get access to like, a new internet connection, I always like, open up speed test and I look at how the latency is, upstream, downstream. And so it's very
Botond Seres 36:35
But FPSes are completely unplayable, over 30 milliseconds. That's what I mean by latency in this case. But you don't get better latency with undersea cables either. Itβs a decent point.
Gene Sigalov 36:47
Yeah. Yeah.
Dave Erickson 36:50
One of the things that I'm I'm kind of curious about a little bit, and we're talking a little bit about it with throttling intelligence on responses, but definitely for say an application like AI receptionists. you know, one of the things that AI, it does in a very strange way. Sometimes it can be overly human, particularly when it's over complimenting people. It loves to compliment people and tell you how great you are. But then sometimes you're, you feel like it doesn't understand your emotions, that it doesn't understand the emotional aspect of what you're doing. In speech to text, text to speech, how do you kind of get it so that the responses of the AI receptionists show that there's it cares or understands the emotional side of what someone's saying?
Sometimes receptionists, they get they get an emotional call, right? Someone's upset, something's not working. And how do you get it so the AI kind of responds in a way where the person on the other side thinks it cares?
Gene Sigalov 38:00
Sever, several thoughts come to mind. So, first of all, obviously prompt engineering, right, where you want it to be kind, respectful, thoughtful, but not overly effusive, right? I, I will say that some I'll answer the question in a slightly different way, a lot of the time, prospective customers worry that the AI won't do very well with the elderly, but it actually does very well with the elderly because it's infinitely patient. Right. It will repeat itself. It'll enunciate things for you. It'll it'll just it has all the patience in the world, right? Which you need sometimes when dealing with a elderly person, and which actual humans don't always have the patience to, to, to treat with the respect and kindness that the elderly deserve. so it's actually a strong suit of the AI that's been properly prompted to engage with people in a kind and thoughtful way that's not overly effusive. But the the models do have their own idiosyncrasies, right? So I'm a pretty heavy Claude user and I'm a pretty heavy ChatGPT user, Codex user, if you will. And I do see that their behaviors are dramatically different in how they engage with people and how they engage with people in or try to engage with people in an emotional way. And I expect that that's just their kind of default prompting, which you layer on top of when you do your own prompting.
Dave Erickson 39:29
Just out of curiosity, between Claude and Chat GPT, which do you think is more I, I guess human centric in its communication?
Gene Sigalov 39:40
I th, I think that one of the observations that has been broadly made by people probably over the past maybe three to six months is that Claude has been better to work with, right? It that, the the the intelligence has not necessarily always been better and certainly not now, right? Like the Frontier models with you know, Seoul five point six at you know highest level of intelligence versus Fable, they're pretty comparable. And I think that people's commentary is that Claude has been more interesting to work with in the way in the human way that it kind of replies. So I think that Claude is probably the crowd favorite. But my personal opinion is that the current state of the art, Chat GPT slash Codex is a better value. I think their app has gotten much better, the way that you can run codex on your desktop or laptop and use it remotely,. just the app they're they've really gotten their act together, right? As a, as a personal observation. I you know, in my personal opinion I ChatGPT is a better value today than, than Claude. Although I continue to use both of them because I don't know what's tomorrow, right? It's, they leapfrog each other pretty regularly.
Dave Erickson 40:54
Then that kind of brings us a little bit more towards application. Now that you've kind of applied speech to text, text to speech to a business model of receptionists. Yeah. And you kind of have an understanding of how it functions in a, a business cycle, what do you think are gonna be some of the, the next real beneficial applications of it? What, what businesses do you think are gonna be able to benefit from these technologies next?
Gene Sigalov 41:26
This is like a blue sky question, like bro broadly, you know, sure broad, broad thoughts. It's a it's a it's a good
Dave Erickson 41:33
Receptionist is kind of easy 'cause it's human to human face interface, right? That, that's a great application for text to speech, speech text. But what are some others do you think
Gene Sigalov 41:45
I honestly don't know, but one of the things that strikes me is that if I were starting a business and I was a little bit younger, maybe I was in my mid twenties, I would do something that's hardware based, and also has a strong software element to it, right? Because software is to some degree not completely, not fully, not production based systems, but in, in, in some I don't want to say primitive, but in, in, in some capacity has been solved, right? And, and, it allows you to really be very creative with new products, physical products. So I, I think that you know, you can get like a Raspberry Pi and you could automate anything in your house, right? Like you want like a I don't know, a geometric array of like, sprinklers to like do like you know some sort of like to to like music, right? Like you can do that device, right? You can get like a Raspberry Pi and you can do like a water fountain display that you know does like a parametric EQ, right? To like the sound of music, right, on your lawn, right? Like you could do that sort of thing, right? Which you couldn't do before, right? Because it was cost prohibitive, right? Maybe you could do it as like you know a six month project for fun, but now you can probably do it as a weekend project, right? To have like a water fountain on your lawn that like plays to like, Lady Gaga or something.
Dave Erickson 43:09
One of the other things is about speech to text is accuracy. I mean, I've noticed it because we've been transcribing this podcast for five years now. Mm-hmm. And five years ago when we would run these into transcription programs and there weren't that many of them, you know, we would get like eighty percent accuracy. And we would go through it and have to correct, you know, twenty percent of it. And you know, and that's kind of moved up and now it's getting closer to like ninety, ninety five percent.
Gene Sigalov 43:41
I think it's like ninety three I think it's like ninety three percent, ninety four, I think that
Dave Erickson 43:46
Yeah, it's pretty much somewhere in there. Do you see it ever getting like to ninety nine or a hundred percent in accuracy? And
Gene Sigalov 43:54
I think the question is, is, what do you think human accuracy is? Like how, how good is your, is your meat processor in between your ears? Is it a hundred percent? I don't know. My wife said Same same year.
Dave Erickson 44:02
It's about thirty percent.
Botond Seres 44:07
Do you watch movies with subtitles?
Gene Sigalov 44:11
sometimes. Some sometimes and and some and sometimes even in English I watch movies with subtitles just 'cause it's it's fun to to kind of turn off the sap and just watch it quiet. yeah, I do.
Dave Erickson 44:24
Yeah, I, I think the accuracy thing. I mean I, and I think part of that accuracy thing is an emotional content because when I'm looking through some of this stuff, some of the inaccuracies are because it, it didn't recognize how someone was saying it emotionally, right? Because you can say something definitely in English, but in other languages also, where you can say something, and if you say it in you know, put the emphasis on the beginning of the sentence, it actually has a different meaning than if you put the emphasis on the end of the sentence. Yep. I think AI is still struggling to deal with some of that. Yep. Do you think it'll get better or what do you think?
Gene Sigalov 45:03
I, I do think it'll get better. and also kind of an unmentioned ac accuracy point. Like we, we here right now we have like microphones, right? Like you'd see b Baton's microphone, right? He even has like the wind thing on it or whatever, right? Like I have a a nice microphone. And so like this is the best environment for AI. This is where you can get ninety-three, ninety-four percent, right? if you're taking like a call in like a windy area, right, like you're in a busy office, right? You know, you have to do noise cancellation and some words won't be caught, right? So in imperfect environments, you don't even get the ninety-three, ninety-four percent, right? That's only in kind of like a synthetic, perfect environment. and these things are getting better, right? Like noise cancellation technologies are improving, right? We use noise cancellation on our phone calls, and we're always looking for better tools to use. Telephony quality is also important, right? Like what's the codec on your phone provider? What's the codec between the phones that are connecting, right? Can you get that high bit rate phone call? So all these things matter and then they all inform the, the quality of the transcription.
Dave Erickson 46:10
When you were putting together Up First and you did the first release, did you what, what kind of surprised you about it, or even just building it from a business aspect, what what kind of surprised you about it?
Gene Sigalov 46:27
What surprised me? I don't think I I nothing really surprised me about the experience of building a product, right? Because it's just another one of those, right? This is like a a proven playbook that I've run before. but what surprised me is that the way that people responded to it and that actually people had come to us complaining about their human services, right? Because the question was always like, is this as good as a human? And in some respect you could say, No, it's better than the human, because it has all the context and knows all the things about your business. Whereas a human, a fractional employee, most of the time, right, these services which we're competing with are fractional employee services where they answer the phone for many, many different businesses. They know very little, right? And so in some respects it's better, right? And again, with that case where I explained that an elderly person, he has a counterpart that's infinitely patient with him or her. And so that's also important, I think, as well. But in terms of like building the actual product, I don't think anything surprised me because you know, so surprises are the rule when, when you're building your own business, as you know, right? the the thing that would have been surprising if there's if there weren't any surprises, that would have been surprising.
Botond Seres 47:54
Yeah, Gene, I was wondering if you do get negative feedback and if you do, what do people not like about the AI receptionist?
Gene Sigalov 48:07
Latency is a, not an, not a frequent complaint, but occasionally people, you know, we have like a little button where people can submit feedback on every single phone call, right? And so, you know, there's kind of like two categories of feedback. There's the like, hey, this service is not good for me, right? So I'm gonna go use a human or I'm gonna go use a different service. And then there's the category of observations that people make with respect to individual phone calls. And, you know broadly, they're kind of the same thing, right? either b bothers them enough to leave or they've just made an observation and hoping you improve on it and doesn't bother them enough to leave. and so I think that occasionally, not understanding a person, like what they're saying, right? Contextually, like being out of sync with the conversation, happens rarely, but it does happen. Latency is kind of like, a, a general complaint, and you know, noise cancellation is something that people would want more of.And we're using the very best noise cancellation technology available today. There is a new technology that we're gonna try and maybe it'll be a little bit better, you know, a, a few more points better. we don't know yet.
Botond Seres 49:14
noise cancellation if the cooler's side is noisy, so they...
Gene Sigalov 49:17
Yeah, yeah, yeah. Because yeah, yeah, yeah. Because a lot of time callers are calling from an office or they're calling on the go, there's wind in the background, there's a dog barking, there's traffic, there's cars, honking, whatever you can think of, you know, not every phone call is on this like kind of like synthetic perfect space.
Botond Seres 49:33
Because our brains are pretty good at noise cancelling. The AI is not as good.
Gene Sigalov 49:38
Yeah, it yeah, the A, the AI is getting everything. Yeah.
Dave Erickson 49:41
Okay. What do, what do you use for noise cancelling?
Gene Sigalov 49:43
We, we use Krisp, K R I S P which is kind of Yeah. So Crisp Crisp is kind of the standard. there are upstarts which are advertising, right? Better performance. You know, there's advertising and there's the truth, right? And the only way to figure it out is to spend the time to implement and to test and to listen and to look at the transcripts. And so we're in the process.
Botond Seres 50:08
And I'm wondering about latency. Do you get the same complaint over and over or do you get both sides? Like is it always, no, the AI is talking over me? Sometimes they ask to...
Gene Sigalov 50:19
Sometimes it's too fast. Sometimes it's too fast. But, but that, that's in the minority. That's a minority opinion. I have heard that, but it's you know, the bigger complaint is it's too slow, right? it's either kind of like the piece Yeah, it's it's it's either the pieces in the stack are too slow, right? Like maybe it's like half a millisecond too slow, slower than they'd like on average, or sometimes it just needed to think longer, or maybe open AI had some downtime and it was just, you know.
they you know, 'cause you you you're you're dependent on the vendors, right? Whether it's OpenAI, it's Grok, whatever it is you're using, right? Sometimes, you know, things aren't peachy. You know, they're having a rough day, right? They have downtime too. Yes, that stuff happens. So th there's kind of like, you know, there there's two categories of complaints, like, hey, this was slow in this particular phone call, or kind of like an ambient slow, right? Like this is always just a a hair too slow for me.
Botond Seres 51:10
But the complaint in general is that it's too slow.
Gene Sigalov 51:16
yes. I also don't want to overstate it. It's, it's not, it's not like this is all our complaints. This is this is just one of the I think things that people in the industry knows needs to be improved when the different providers who are part of the stack, you know, whether it's open AI or it's 11 Labs or it's Deepgram or whatever, you know, Cartesia, you know, there there's inworld, there's there's so many vendors, they all advertise, you know, how many milliseconds, right? How many milliseconds until you first voice, how many milliseconds you know, on average. so everyone's very conscious of latency and they're trying to drive latency down, right? Something always on the mind.
Dave Erickson 51:57
The, the application of AI receptionists that would be most interesting in a sense to me, and I don't know how AI how, how it works or how it would work is a spam filter. A lot of people, if I put a phone number up publicly on a website or whatever, a lot of the calls are going to be spam, people trying to sell you something.
Gene Sigalov 52:18
Yeah, we, we drop spam calls. Yeah, we, we determine that contextually, right, based on the content of the conversation, right? So it's like semantic spam detection. And after about thirty seconds or forty seconds, when, when the AI understands this is a commercial solicitation, it drops the call and it doesn't charge the user for the call. So that's a free call. That's on us.
Dave Erickson 52:38
But what if the, the spam is just somebody calling in? Can you does the AI in kind of in a sense interrogate them to trying to find out why are they calling and what are they trying to
Gene Sigalov 52:49
Yeah, it just handles the phone call like it would handle a regular phone call and one, once it 'cause it it's not a human, obviously, but it has that kind of like human level of intelligence and once it understands it like, Hey, these people want my money, it just drops the call.
Dave Erickson 53:03
Got it. All right. Cool.
Botond Seres 53:06
Jean, usually my last question is about the future. In this case, we can either explore the future of AI in general or the future of AI speech. The floor is yours, whichever you'd like to talk about.
Gene Sigalov 53:23
I think AI in general is much more interesting than my, my little cubbyhole with AI speech. I think it's what we're all to some degree enthralled to at this point in time in 2026 and there are much more existential questions of interest than my little neck of the woods.
Botond Seres 53:43
Alright, so in your opinion, what's the future of AI? Go as far as you'd like.
Gene Sigalov 53:51
It's a, it's a, it's a great and interesting question. And I think that the best way to predict the future is both to extrapolate kind of the step change that we've experienced in the last three years and to be mindful of great thinkers like Ray Kurzweil, Kurzweil, who says that people lack an appreciation for the exponential nature of technology, right? Because people are built to think in linear terms naturally. It's natural to think one, two, three, four. But what we're experiencing is a step change in performance of these models that's difficult to under, to overstate. Right. If you think about where we are just with the coding, for example, some my team, the developers, were screaming that the LLMs suck and they need huge oversight as recently as November, December 2025. And here we are in August 2026. And it's taken for granted that you must use LLMs for coding and that they do most of the work, right? And that's just like six, seven, eight months, right? So the amount of change that we've experienced in less than a year is so dramatic that it's difficult to overstate. It's like, why isn't everyone like screaming their heads off about how crazy this is? Right. It's absolutely insane. And so I think that trying to extrapolate into the future is both very difficult considering how fast things are changing. And I would leave it to the great minds like Ray Kurzweil to to to inform us on how the next 10 years are gonna go,
Right. And I would encourage everyone to, you know, read the singularity or the si singularity is here. he's really the best person to inform us of the playbook for how things are gonna go. Right? He's predicted things 30 years ago, which everyone thought he was insane for predicting. If you don't know who Ray Kurzweil is, you should know who he is, you know, he's child prodigy. He's credited with inventing optical character recognition. He's credited with inventing a synthesizer. I don't know if it was the first synthesizer, synth synthesizer, but he's been most accurate in predicting the technological curve of the last you know 30, 40 years. And I think we'd all be well heeded to to look at his work to understand how things are, are likely to go in the next little while my my w which is which is all a long way to say. I'm not in the predicting game, right? He who lives by the crystal ball eventually, you know, eats shards of glass. And so I I I I would defer to, to greater thinkers than I.
Dave Erickson 56:48
All right, well Gene, tell us a little bit where you're going with up first and a little bit more about what your, your future plans are.
Gene Sigalov 56:57
Yeah. So with regard, with regard to Up First, we're looking to really fill out that AI communication offering for small businesses. So not just reception service, but to handle outbound calls, to handle inbound calls, to handle text conversations, to really be a one-stop application suite for small businesses for their communication needs. So they can place phone calls, they can send text, they can have the AI place phone calls, they can have the AI reply to text, and all for a low cost, right? Without increasing what we're offering, just providing a richer offer, richer offering at the same price level. and once, once we, we feel like we have a a good horizontal feature set kind of in the, in the telecom domain, I think we're going to focus or pay some attention specifically to attorneys. I think that's our, our most natural niche to go after because I'm an attorney. I know many attorneys. I know I know their pain points. And I think more and deeper integrations with attorney CRMs like Clio, building out more workflows for attorneys. So I think that's the path that we're on for the time being. Of course always subject to change as new information comes along.
Dave Erickson 58:18
And just out of curiosity, I assume that your systems are multilingual, is that correct?
Gene Sigalov 58:23
Yeah, yeah, yeah, yeah, yeah, yeah, yeah. I I I I don't think they're quite as good in all the languages, but they're very strong in the main languages, right? English and Spanish. I think it depends a lot on the corpus of training that OpenAI, for example, had, right? English is obviously lingua franca, right? So their biggest corpus of training is in English, Spanish has a lot, but you know, obscure dialects, obscure languages, right? You know, it, it supports maybe seventy, seventy five languages right now. but they're not, it's not as strong in all the languages, right? It's probably not gonna be as strong in Hebrew, for example. Mm-hmm.
Dave Erickson 59:01
Well, Gene, thank you so much for being on our podcast and discussing AI speech technologies. We are at the end of the episode today, but before you all go, we want you to think about this important question.
Botond Seres 59:15
How will you change your business communications with AI? For our listeners, please subscribe and click the notifications to join us for our next ScreamingBox Technology and Business Rundown podcast. And until then, see if you can get your business AI to talk to your customers.
Dave Erickson 59:34
Thank you very much for taking this journey with us. Join us for our next exciting exploration of technology and business in the first week of every month. Please help us by subscribing, liking, and following us on whichever platform you're listening to or watching us on. We hope you enjoyed this podcast and please let us know any subjects or topics you would like us to discuss in our next podcast by leaving a message for us in the comment sections or sending us a Twitter DM. Till next month, please stay happy and healthy.
Creators and Guests
