Competition: The Sentence II

With ‘clipping’ do you mean the crackle that seems to pop up around consonants

The crackle, plus the sounds don’t “flow” into each other, they sound like they are just strung together with tiny gaps in between. It’s hard to describe. In fact it does sound like reversed speech in places. Imma play it backwards to see if Cognis is actually dead and has been replaced by a lorry driver from Brighton.

  1. Hen hongu evie.
  2. Ang eel ronk ang wast it in the head an a(ch)uruk, among a mar hare.
  3. Here, I’m at our uncharade.
  4. What pull ill wheel, where are we?
  5. You’re in some earmott. (or earmah?)
  6. What? Ong uh that Your companion was cased in Iraq.
  7. Some earmott is technically in Iraq.
  8. It’s technically ouah and we’re in the middle of it.
  9. Is ersnot. Ersnot is in ra-eeing. Kumpah!
  10. Love in error. Poop. Loves in the S’Kong. Ang eel na soviet il’yangeles dobu (beep) up un ist kaos erkay ayu est.
  11. Kung Ishing, honor ronk is ready.
  12. What wrong, an ung it an et mar enuen.

Ow, my brain.

i say nebular wins :wink:

Umm…

1: Hey how are you feeling?
3: Yeah, I had a rough trip.
4: What do you mean? Where are we?

That’s about the extent of what isn’t unintelligible. I think I want this program…

I’d recommend Festival. I was playing around with it on my computer and it sounds, well, like English. Not… possibly obscene.

Well… have a good one!

No, never knew it had one! I have a feeling it would be proprietary software, though, making it impossible for us to expand it. If not… :smiley:

I think I know what you mean, and this might just be something I can adjust directly. I will try, and upload the adjusted version some time this weekend!

<Grabs Dark Talisman>Kong here am ah, Kong here am ah…

Well, you can get the program, no problem. It will be made Open Source at the end of the competition, in the hopes that we can squash a few more bugs out of it with the help of all your feedback.
As for Festival, it’s already being looked at (thanks for suggesting it though; we are very open to outside ideas at this point :D). It is a question of how much is proprietary (sometimes the software is Open Source, but voices are not, f.ex. AT&T’s Natural Voices used in an OS program) and what the options for further custom features are. And, of course, quality. As previously mentioned, MARY was considered, but it’s good voices seem to be proprietary, and the other voices are problematic and aimed more at German :frowning:

I began to ponder on this… are you saying that when finished this is going to sound better than AT&T Labs Text-to-Speech (Try it here: http://public.research.att.com/~ttsweb/tts/demo.php) for example? I couldn’t get MARY to work (some server errors on their side) but I haven’t heard any better TTS than AT&T Labs demo.

Also, if you are indeed somehow going to get this to work in this year with same level of quality, that would be awesome :slight_smile:

Is this going to be proprietary program or what kind?

Opera speak sounds really mechanical also compared to AT&T Labs demo… try it out and you’ll be amazed :wink:

EDIT: Oh you beated me to tell that this is going to be open source and that you already know AT&T…

i think it’s absolutely amazing that you are working on a speech synthesis engine, but i’m curious as to why you say that MARY is the only open source engine around.
there are quite a few listed here:
http://linux-sound.org/speech.html

We hope to have something comparable to AT&T for basic voice generation before the release of Siberia Complex later this year, but it is a hope; whether we succeed is up to us. Also, AT&T is IMHO not the best on the market, check this. Ours will be far more limited and less user-friendly than any commercial one, but the first objective is to simply get it working, and making it understandable (and not painful to hear:o).
And yes, it is the plan to make it either entirely or partially Open Source. There is a possible offer, if we achieve a sufficient quality, but it would only be a form of license, making the engine OS later or as a “less pretty” version (Open Sourcers like it a bit rough, after all ;)).

Well, I do know that there are others; Festival/MBROLA is being looked at at the moment. MARY is just the only one we have been able to actually get working as “advertised” on our own systems :rolleyes: I looked through the list, and there are many interesting ones I have never heard of before, though! Come monday, I will be pushing for a check of as many as possible.
The problem with many of the Open Source speech engines is that licenses are often murky, because the voice libraries (little voice samples used to construct a voice) are sometimes proprietary, even if the program is not. Also, many OS speech engines are mainly just academic experiments, which get half-finished and poorly documented, making it tough to continue the work. There are high hopes that Festival/MBROLA can come to our rescue, should our own speech engine hit a wall, or maybe Festival/MBROLA will just turn out to be perfect for our needs.

Just wanted to make it known that I received an update this evening (apparently audio people work weekends:eek:) and I wanted to post it. It is not as much an improved version as an alternate approach; we rolled back some functions to get rid of a pesky problem with consonants. If people could compare how easy it is to handle as opposed to the original dialog sound files, it would be great. You don’t have to in order to fight for the prize, though :wink:

Download RAR pack

And I was asked to declare openly that the speech engine has now passed 200 man-hours of work! Hope it’s worth it :slight_smile:

There are a lot of incoming guesses, and while no 100% match has come in, people are getting really close to it.

Because of this, the competition has now been given a deadline:

Monday April 30!

Unless a full 100% hit is submitted before that day, on that day a winner will be announced by “Closest submission”.

So if you did not submit your guess by PM or posting yet, now is the time :eyebrowlift: !

Okay, I’ll have another go:

  1. Hey, how’re you feeling?
  2. I feel like I was hit in the head by a truck! Oh my god, your hair!
  3. Yeah, I had a rough trip
  4. What do you mean? Where are we?
  5. You’re in Siberia
  6. What? I thought that your company was based in Europe
  7. Siberia is technically in Europe
  8. It’s technically nowhere. And we’re in the middle of it
  9. It’s Russia. Russia is European. Come on!
  10. Whatever! Dude, what’s going on? I knew the Soviet Union was deserted, but this is just ridiculous!
  11. Quit bitching! Our ride is waiting
  12. What ride? And how did I get here anyway?

It’s an improvement I think, the harder lines are a lot easier to work out now:yes:
I’m not sure what “ride” was in lines 11+12, I just guessed that one…:wink:

Cognis,
I think you should have called it Bleeding Ears :slight_smile: OK. I really could not make any confident guesses, so I’ve typed each one phonetically to tell you wht sounds I think I heard. Rather than guessing what it could mean, in most cases, since I’m pretty sure that did not come across. (sorry)

  1. Ping Pong on TV.
  2. An e ran kon un see ima eh kann irhe, an ran ka more air.
    (was this in German?)
  3. ihr an herr I ah ree
  4. What do you mean where are we?
  5. You’re in con here not. (‘Con here not’ probably one word? a place?)
  6. What on earth have you compan ___ ____ ___ in eraq.
  7. I’m here not in ___ ___ ___ in eraq.
  8. It’s definately over there and we’re in the (name) of it.
  9. I’s r’not, r’not its ering can ta
  10. Ur ehnr, oo, oo est in est con …
    (remainder too fast I have no clue.)
  11. Who eh, is she, our {car} is ready
    Not sure about ‘car’
  12. What e’rang an kam eh a enyah en rate.

Well, I hope this helps.

Regards,
Mike

It’s great to see the submissions still coming in. Some are almost there, and those that are not are providing a LOT of information useful to our work (we’re not doing this competition to brag, we’re doing it because you point out our flaws very nicely ;)).

I just wanted to note that the full C++ code for the speech engine is now available here. It does not include the lines used in this competition; those will be released in phonetic writing for the engine once the winner has been found. documentation is also very limited at the moment; in particular, the phonetic ‘alphabet’ used needs to be described clearly for anyone using it. The few examples included give a starting point, though.

So now all you code geeks can play around with it, too! Read the posts in the thread, though; they contain some important information!

  1. I have to say that I believe the movie White Noise was NOT fiction :slight_smile: Seriously though, I do think that if you listen to anything long enough and let your imagination run wild, you will hear things; your brain is wired to do that.
  2. Perhaps it is hard to understand becuase Cognis is embedding subliminal cues to dontate money to his paypal account.
  3. So how does a “kong here am up” become “Siberia”?
  4. Cognis: yeah, you’re right about the inflection and stuff, although you can raise pitch just by speeding up the bitrate, so like to use a word at the end of a sentence just accellerate its speed. I realy did think that 99% of words used in common commuincation were a very small set.
  5. Have you looked at the way the Microsoft Windows speech synthesizer works, because it is readily understandable (ignore this if you have already answered this in other threads). :frowning:

:eyebrowlift:

Knowing nothing whatsoever about speech synthesis I’ll offer a little advice :slight_smile:

It seems to me from listening to a few of these that the engine needs to operate on sounds, not letters. To me, currently, it just seems to spit out every letter fed to it so “cat” becomes “KKkk - a - TTtt” where the two consonants are delivered as if a child spelling the word, phonetically, out loud. I think this is what CD38 is referring to. Where you’re getting better results is where the consonants in the word have softer phonemes that blend nicely into the vowels (like “where”).

I liken this to lip-syncing where good results depend on animating actual sounds, not the letters that happen to make up the spoken words. This is why I have reservations about the usefulness, currently, of automated lip-syncing solutions. Letters become effectively redundant in many spoken words. For example, in the word “don’t” we can’t really see the “n” in a viseme and in reality, it barely registers as an individual phoneme since it combines the the “t” to make a unique “nt” sound, You engine however, would currently deliver the word as a rapid “Duh - OH - uN - TTtt”, turning a one syllable word into four syllables. The word which convinced me of the problem is “technically” (assuming that is the word in #8). I didn’t guess that word but when I saw others had, I listened again and noted the individual “TTtt” and KKkk" sounds that added extra syllables. In reality, you’d barely notice the “ch” sounds in that word and the “t” would be glued to the following “e”.

So, for voice-synth I’d imagine a huge library of common sounds like “ta, tA, te, TE, to, TOO, etc…” where the lowercase letters show short vowels and uppercase show long vowels. This means “cat” would be spoken using either “ka - t” or, for even better results, “cat” where these phonemes hare basically hardwired.

Of course, this would ultimately mean that the text be either typed phonetically with regard to this phoneme library or the engine have a dictionary in which all words have already been broken down into these phonemes. There can only be a few million phonemes, can’t there? :wink:

You’ve probably already considered all this stuff but I figured it’s worth mentioning since you say your wondering what makes some things work better than others.

I applaud your efforts and look forward to progress.

but andy, that would mean having a dictionary that translates “don’t” to english “DUH-O-n-T” and german “vas nee-sh-t” or (whatever it is in german) etc where the translation is a standard common sound the engine puts out. And for a general purpose, it would mean typing in an entire dictionary. I think Cognis tried this using standard dictionary notation and it failed. The only other idea, to build on yours, is to construct a pronunciation guide that algoritmically translates a written word to a spoken one using a set of rules.

For anybody who’s interested… a PDF of an old paper I read in college:

http://ieeexplore.ieee.org/iel6/29/26121/01162873.pdf?arnumber=1162873

Letter-to-Sound Rules for Automatic Translation of
English Text to Phonetics

HONEY S. ELOVITZ, RODNEY JOHNSON, ASTRID McHUGH, AND JOHN E. SHORE, MEMBER, IEEE

Sarah

(the speech engine is from here on called ‘Chad’)

Andy, you are only part correct, but in that part, you are completely right :slight_smile: There are groups of ‘voice sounds’, which we sort around. We sort them into soft, hard, and sharp sounds (in the Chad code commentary, ‘sharp’ and ‘hard’ sounds are not clearly distincted between, so ignore that for a second if you read it).

Soft sounds include all vowel, and consonants that activate the vocal cords; m, n, r, and l are examples. Chad stores all of these as ‘waves’, since speech containing these sounds simply repeats such a wave ‘print’ over and over again to produce the sounds (in reality, it alters the details of the wave, but the exact method still eludes us, as it is small details. We use pitch and ‘bend’ to emulate it; without that, the voice sounds very metallic). Soft sounds are mixed by Chad, so that the resulting sound phases smoothly between each soft sound; ‘l’ phases smoothly into ‘a’, which phases to ‘i’ which phases to ‘m’, to create the word “lame”, for example (sorry, only soft-sounds-only word I could think of without warning:o).

Hard sounds are sounds like s, h and f, which do not invoke the vocal cords but can be maintained; you can say the ‘s’ sound for as long as you have air. These sounds do NOT use waves, but rather different forms of random white noise patterns. We actually reproduced the ‘f’ sound with pure math, and are working on others. Chad simply copies hard sounds directly into the sound file.

Sharp sounds come in a snap and cannot be maintained; p, d, b, k, and t, for example. Chad simply copies these into the sound file, too.

So yes, “cat” is ‘just’ placing a ‘k’,‘a’ and ‘t’ sound together, but that is because this word has two sharp sounds and just one soft. “Kate”, for example, would be ‘k’, then an ‘a’ phasing into an ‘i’, and finally a ‘t’ just put in (actually, both cat and Kate would use a ‘d’ rather than ‘t’ due to the pronounciation, but this is not a lecture, so I take shortcuts :p).

We do have problems with it, as you mention. The mouth has a split second in which it produces a ‘mid sound’ between any hard/sharp sound and a following soft sound (it is less common if the soft sound is first, which might explain so many correct guesses on “middle OF IT”, #8). Thus, words like “buy”, “do” or, yes, “Siberia” sound ‘chopped’ or ‘clipped’. At the moment, there seems to be 2 solutions to that: Add such mid-sounds to the sound library, or find a way to calculate them during processing. Either requires a lot of work, and neither is guaranteed to work :mad: Nonetheless, we’re working on it, and any suggestion is welcome.

What seems to also be a strong factor is the aforementioned miniscule variations made to especially soft sounds in speech. These variations keep our voices from creating harmonic vibrations, which (in spite of the blissful term) are a living hell. At best, they make voices sound “robotic”. At worst, you get high-pitch whines for brief glimpses. Those whines are bad if you are listening through ear phones, I can tell you :frowning: Chad uses variations in pitch with every copy of the soft sound waveprint inserted to reduce it, as well as ‘bend’ (a mathematical alteration of the soundwave, too complex for a brief explanation), but it is not enough (and sometimes creates extremely ‘raspy’ voices, like extreme smokers have. Or bad horror movie demons).

We have some research results on these variations which will be put to use in the next version of Chad. If they work, we dance, if not, we cry…

Keep your thoughts coming on this, we have hefty debates on all of it, even if you go on a wild hunt (perhaps especially if that happens…). Ideas come from people speaking their minds, not from people waiting till they understand everything (do we seem like we’re the masters of this game? :p).

@Roger: Chad uses a phonetic writing system; words are written as their sounds, not their proper English spelling. This can grow on its own into a dictionary, once Chad works as it is supposed to (i.e. if you want a word not in the dictionary, you can enter it, just like on many cell phone text messaging systems). One advantage of this is that you can manually alter the pronounciation of a word to fit the given circumstance, like how “how you” is merged in #1. the major drawback, of course, is the need to write every word correctly, which requires some serious linguistic skill, and often multiple attempts; not good for efficiency. A design to get the best of all is underways. Suggestions are welcome.

@SARAH: I already took a look at the article. Do I need a login? If not, or if you can provide me with info on getting one, I will look at it personally, ASAP.

PS: Two participatants are very, very close to a 100% correct submission. I will not tell you who, though (and remember, I get submissions by PM, too ;))

Just wanted to remind people that there are 5 days left before The Sentence II closes for submissions. There have been far fewer submissions these last few days, but in accordance with the April 30. deadline I will still allow these last few days for those who forgot to submit, or did not yet notice the competition.