Competition: The Sentence II

EDIT: It should be noted that if this competition proves our speech engine to be of at least some level of quality (judged by responses and the amount of lawsuits due to hearing loss), its current incarnation will be made Open Source, while we work on a more advanced version.

As some already know, Siberia Complex is a movie underways using Blender for most visual aspects. As some also know, we’re working on how far computer speech can be pushed to be worth using in a movie. While there are plenty of commercial speech engines around, they allow very little to no tweaking or code improvement, and they do not sound all that good. The one Open Source engine around (MARY) is still a work-in-progress, that also uses several proprietary voice components and is actually best used for German speech.

So we are working on our own.

The first The Sentence contest in these forums had just one line; now we are upping the ante to a 12-line dialog! The Sentence II is being held to see how much work must still be done in improving the voice outputs. We already know that it sounds horrible (think computer voices from back in the 80s), but is it at least understandable?

To find out, this competition offers the winner a free pick at the Blender Eshop (or a comparable gift certificate to Amazon. Sadly, gift certificates are apparently not available for the Blender-shop:(). All you need to do is listen to the dialog and post what you think they are saying in the dialog. Please write your answer in the following format:

Person 1: Blah blah blah
Person 2: Blablabla
Person 1: Blah? Blah blah blabla

…and so on. Your answer must include all 12 lines of dialog, but anyone may post as many answers as they like (please post one full dialog only per post, or something might get overlooked).

The winner is whoever gets the dialog sufficiently right. ‘Sufficiently’ is determined by the people trying to make the speech engine, but will be based on about 90% correctness of words written (the first contest had a winner writing “how’re you feeling” rather than “how you feeling”, for example).

Occassionally, new and (hopefully) improved versions of the dialog will be uploaded, until someone guesses it. We are still working on making the voices more pleasing to the ear (a subtitle for the competition was considered to be “The Sentence II: Bleeding Ears”, but it seemed overly dramatic :)).

The dialog is available for download here. It contains all 12 lines as wav files, in numeric order. The voices are still fairly similar, but one person is slightly higher pitched than the other.

Good luck!

You want us to post our answers here or PM you? I think it would skew the results if we saw each other’s responses.

WinRAR gives me error messages when trying to extract the .wav files.
Is this normal?

So these are going to be the voices for your movie?

Either would be fine. Of course, there is street cred for daring to show your skills in public :cool:

EDIT: When the competition is over, we may choose to do a “Hall of Cool” for the guesses that are simply so funny or original that we feel like sharing them with others, like RogerWickes excellent " Hey, dam your spewing!" guess in The Sentence I. We’ll ask permission (and offer anonymity :D) at that point, though.

No, it should not be. It was ZIPped with WinRAR, so it should open easy. I’ll have a look at it, and put up an alternate if I find something. Expect a result later today or tommorrow.

EDIT: Seemed like a good idea to make sure quick, so here is an alternate RAR package. For ‘emergencies’ :confused:, you kan download (or listen to) each file directly: 1 2 3 4 5 6 7 8 9 10 11 12

LOL, hopefully not! Note that I described them as ‘horrible’ (‘Bleeding Ears’ was a serious consideration:eek:). The idea is to get feedback from people to determine what makes them sound wrong to others before fixing them entirely. We all worked with them so much we have to continually bring in fresh ears to tell us what is wrong; this is just a way of doing that en masse.

Hopes are that acceptable voices can be made in time for SC’s release (still no date, but 2007 is a serious expectation). If not, we will use computer voices whereever they can do the job well enough, and use real voice actors for the rest (possibly using voice alteration software to get the voices we really want… good voice actors are hideously expensive, especially for lead roles). The philosophy is that if the speech engine can be brought up to the task, it would take less than a day to produce all voices. That, of course, is just theory :confused:

Thanks, it worked this time.
However, what language did you say these people are supposed to speak?
Sounds like “k’tpmwa k’tmwi” to me. I can’t make out any language that I know of (including German and English), though it sounds like chinese or japanese to me.
No, I honestly can’t understand what they’re saying :eek:

It’s English. At the moment, about 15% of testers can make out most of the dialog when listening closely, and another 20% or so get a rough impression. That’s counting the guesses made by people here, plus test subjects IRL. The rest seem to make out bits and pieces, or nothing at all.

The engine needs a lot of work, still, but I am getting some good contenders in PM already, hoping that we can push those 15 and 20% higher :eyebrowlift:

You still have some work to do then. I’ll listen more closely next time I try but it certainly needs work, no offense.

None taken, I am fully aware of that. Thus these competitions :slight_smile:

Hehe… I must be interpreting it so wrong, but here it goes anyways. Overall these sentences leave REALLY much room for a guess… After short listening I got this funny nonsense :stuck_out_tongue:

  1. Hey, how you feeling?
  2. I feel like I almost hit blahblah in error, oh my god my head
  3. Hear, we had a error onroute
  4. What do you feel, where are we?
  5. You are in an blahblah
  6. What, we are that blahblah lost space in error
  7. blahblah cheking in error
  8. It’s checking the area and we are in the middle of it
  9. ?
  10. ?.. lost in US
  11. Who is she? Our error is your doing.
  12. What error and how it blahblah anyway?

EDIT: Oh and I should probably mention that english is not my native language, but I would say that I have at least satisfying understanding of it.

  1. ten, tamoof ewing
  2. omf eong rong kong wass keup eu auer kachera ang wong kah mah ere
  3. ear, ong ad rachewheer
  4. Why pull in weer, where are we?
  5. You’re in Kong here am ah (btw, this one makes a GREAT chant)
  6. what um up lat ere cum pan-u-up est in ere up
  7. kong here am up is check nique in Inraq
  8. Is checknique in erruuh, and we’re in the middle of it.
  9. its reedy ready it’s in rah being, kong pah
  10. poop kong worth repeating that (are you sure this isn’t like, talking backwards?)
  11. kooey fishing, ong wrong is ready
  12. What iran, and kong pi eh pity ong anyway.

I declare myself the winner of the first trial, btw.

btw, you need, in your text annotation markup language, the ability to add pauses for commas, inflection for questions and emphasis, general speed speaking rate, gender, brogue/accent, and emotives for sadness, alarm, disgust, as well as standards for euphemisms, like aw, huh, ewww gross, that sort of utterances.
Like: “Hey, we’re not gettin outa here, are we?” spoken like a true Scottsman. Keep working at it, this round sounds much better than the first contest.

just to be annoying, since it was not Elvis but the Beatles Abbey Road vinyl that had the Paul is Dead backwards, and we should never confuse the King with the Kings, with all the time you’ve put into this, why not just use samples? There’s got to be only a few hundred different vocal sounds, or even a few hundred words that are used in common speech, and it seems it would be a HECK of alot easier on you guyes if you just recorded them and indexed them somehow, and then strung them together. That’s what we did for the first ever NASDAQ automated stock quotation system back in 1984, and it worked pretty well. Just have the vocal talent donate the recording to the public domain. Heck, even some BA-heads would probably send you samples as well if you just gave us a script, sorta like The Quick Brown Fox Jumps Over the Lazy Dog - but the phenome equivalent.

Nice to see some good attempts. However silly some of it may seem, you are actually catching quite a lot of bits and pieces! For some reason, the latter half of 4 (“where are we”) and 8 (“and we’re in the middle of it”) have not been guessed wrong yet, and that includes both submissions here and IRL testers. So we must be doing something right :smiley:

Yes, that is definitely needed. I would like to focus on people understanding the voices first, though :o

  1. Hey, how’re you feeling?
  2. I feel like I was hit in the head by a truck! Oh my god, you’re hair!
  3. Yeah, I had a rough trip
  4. What do you mean? Where are we?
  5. You’re in somewhere, man
  6. What? I thought that your company was based in England
  7. Somewhere is technically in England
  8. It’s technically nowhere. And we’re in the middle of it
  9. It’s Les Sciences. Les Sciences is in bwee-ay. Kong-top!
  10. Le Bwouh. Trouve. Ou est il y a KONG! I feel that serviette il est un pouvou et ist ist est part noust ist ien!
  11. Who is she? Our bond is wordy.
  12. What bond? And how did I get here anyway?

It’s okay up to 9, at which point it sounds like they’re speaking Franglais, then it becomes a jumble of random French words, then at number 11 it becomes fairly understandable again.

ROFL! No french, please dear Lord no French…:smiley:

But the English parts are pretty close, with a few exceptions…

I wanted to be sure I understood what said(/wrote) before I answered. Sorry for that minor delay :wink:

Re: The King(s), seriously, you are not cleared for that, and THEY know where I live, so don’t reveal the Master Plan, just yet :RocknRoll: :cool::evilgrin::ba::RocknRoll: (okay, that is more “Village People” than Beatles, but my smileys are limited…)

Re: The rest: It’s basically a good suggestion, but sadly you make one wrongful assumption: The number of words is not as limited as you might think, especially because the system is being produced for use with a lot more than just Siberia Complex (I am becoming painfully aware that next to no one reads the blog :frowning: , in this case this entry). Not just that, but as you mention, it will at some point be built up to handle emotional, individualized (‘unique’), fluctuating and more voices, not to mention the fact that the framework already works for international languages almost as ‘well’ (/poorly) as it does for English. With just a database of words, it would be easier, but 90% of the future plans for it would be impossible.

The speech engine, as well as the rest of AMPS, are very long-term projects; Siberia Complex is just a field test. We are allowed no shortcuts (or we’d be doing it as a silent movie :p)

Here’s what I have:

1 Hey, how’re you feeling?

2 I feel like I was hit in the head by a truck. Oh my god. Your hair!

3 Yeah. I had a rough trip.

4 What do you mean? Where are we?

5 You’re in Siberia.

6 What? I thought that your company was based in Europe

7 Siberia is technically in Europe

8 It’s technically nowhere and we’re in the middle of it.

9 It’s Russia. Russia is European. Come on.

10 Whatever. So, what’s going on? I knew the Soviet Union was ?building a base in Snosk with the US?

11 ???Who is she? Our ?rock? is working.

*12 What rock? And how did I get here anyway?

Surprising that people aren’t getting Siberia, given the project is called Siberia Complex. I’m finding 10 and 11 almost impossible because of the clipping. Anyway I hope this helps move the process along.

@roger: Revolution 9 from the White Album had the backwards stuff (“numbah nine” becomes “turn me on dead man” if you listen until your ears bleed). But yes, people “discovered” clues all over everything from Sgt Pepper forward :rolleyes: :rolleyes:

With ‘clipping’ do you mean the crackle that seems to pop up around consonants? Or is there something else?

@roger: Revolution 9 from the White Album had the backwards stuff (“numbah nine” becomes “turn me on dead man” if you listen until your ears bleed). But yes, people “discovered” clues all over everything from Sgt Pepper forward :rolleyes: :rolleyes:
You guys are playing with fire, and I am not going to protect you when the Alien Demon Overlords of Reversed Subliminals come to claim your livers :stuck_out_tongue:

  1. Hey, how are you feeling?
  2. mumbles and chokes …hong kong… mumbles and chokes
    3-12. mumbles and chokes
    :stuck_out_tongue:
    It’s how I hear it.

Did you hear the Opera browser’s voice? It’s best I’ve heard so far.