>Good old academic's bait and switch. Get them to pay for the material, >then tell them at the end its useless. :-) > >SteveActually, many if not most of the apps in that book work quite well, and there are a ton of publications since then that prove it. Mark
Separation of speech from multiple speakers speaking in different microphones connected to a mixer
Started by ●May 2, 2008
Reply by ●May 6, 20082008-05-06
Reply by ●May 6, 20082008-05-06
>Can someone give me an answer please? >How many voices or instruments that are singing or playing at one same >moment can some ordinary person distinguish? Assume that person just >listens the recorded mono file and says here are N persons in a choir >or here are M different instruments playing. A person is, say, an >ordinary musician, not Mozart. Also let us assume each voice or each >instrument may play an individual note. >Thanks in advance. >Vladimir. >Sorry, I answered the other part that was "are there any algorithms?" I don't know how many people an individual person can separate. An algorithm, on the other hand, is a different subject. With people, having two ears in a room helps with spatial separation, but once the signal is recorded, i.e., a mono track, spatiality is lost other than perhaps an indication of loudness (variance). Humans have the largest neural net known to exist, btw, and the ICA methods (note the plural) essentially represent neural network processors. They also suffer (often) from scale ambiguity so variance differences between speakers or instruments might not be of help. Frequency (pitch) and formants are the two biggest separators that I can think of in terms of human speech. Mark Mark
Reply by ●May 6, 20082008-05-06
On May 6, 3:26�pm, "markt" <tak...@pericle.com> wrote:> >I do know that people call both the one channel multiple speaker > >problem �(AKA cochannel speakers problem), and the two channel > >multiple speaker problem, and higher numbers of channels multiple > >speaker problem, =93the cocktail party problem.=94 �Which of these do > you > >suppose has any relevance to what happens at a cocktail party? Wait, > >don=92t answer! �It really doesn=92t matter in my response as you will > >see. > >The 3 audio sources > >are indeed linearly summed but not with the same weights, so they > >produce 3 different signals, not all just mixed into one signal. > > Big deal, whether they are scaled or not. �LINEAR means sum plus scale. > This is a basic concept in linear algebra. �That has no bearing on the > cocktail party problem. �The goal of ICA, or any linear separation problem, > is to estimate the mixing matrix, scales included. �Most ICA solutions have > a scale ambiguity anyway (though there are ways to work around any phase > ambiguity). > > >Because of these complications, audio separation is a > >largely unsolved problem.=94 �Wow! �Just as I'd remembered it. So page > >446 contradicts page 147 regarding this being a "practical > >application" of ICA. > > Um, gee, the book was written in 2001, and it is now 2008, hence my > comment "dig deeper." �I never said the book solved the problem. > > > > >Now you have been quite generous with your characterization of me as > >=93ignorant=94. A common definition of "ignorant" is =93lacking > information > >or knowledge=94. > > Yes, clearly. �You ignorantly assumed that I said the book solved the > problem, and you STILL don't understand the difference between a problem > statement and a potential solution, as evidenced by this silly post. �You > also put the book down and assumed that since it had not been solved, it > could not be solved and thus "ICA is no good." �Good detective work there. > > > > >Without a doubt you are =93ignorant=94 of the contents of > >your own recommended reference regarding the =93cocktail party > problem=94. > >Apparently my memory at my almost AARP age still functions better than > >your own. > > I fully know the contents of the book, and not once did I ever say that > the book solved the problem. �You did, not me. �Mark, You are fabricating again. I never said the book solved the problem. Pay attention to the details. And deal with those poor self-esteem issues :-). It takes a man to say he's wrong and move on. Then it's no big deal. Everyone is wrong sometimes. It's been a pleasure, Dirk simply provided that as a> reference. > > > rant > > You're still ignorant. > > Mark
Reply by ●May 6, 20082008-05-06
>Pick any research topic and you will get a flood of hits. You'll even >find plenty that claim wonderful things. Do you know of a hit backed up >with a wave file or two that sounds impressive?No, I merely provided a starting point for investigation. The investigation itself is not my job. I did run across a dissertation yesterday while I was cleaning up my references for another work and they seem to have a solution. I did not read it since I don't really care as I typically do radar/comm. ICA really hasn't gotten much attention till the past few years. Mark
Reply by ●May 6, 20082008-05-06
>It takes a man to >say he's wrong and move on. Then it's no big deal. Everyone is wrong >sometimes.So, what exactly am I wrong about? Oh, that you didn't say I claimed ICA had solved the problem in the book... so what's your point about all this ranting then? Curious, since I never made any claims about it's validity. So, we're waiting, go ahead. While you're at it, look up the term "linear" with respect to mixtures. I'm still laughing at that "point" you made. Mark
Reply by ●May 6, 20082008-05-06
Vladimir Malakhov wrote:> On 2 май, 20:00, "markt" <tak...@pericle.com> wrote: >> This is the well-known "cocktail party" problem. > > Sorry for coming back to initial problem. > Can someone give me an answer please? > How many voices or instruments that are singing or playing at one same > moment can some ordinary person distinguish? Assume that person just > listens the recorded mono file and says here are N persons in a choir > or here are M different instruments playing. A person is, say, an > ordinary musician, not Mozart. Also let us assume each voice or each > instrument may play an individual note. > Thanks in advance. > Vladimir.I don't think that anyone can tell how many second violins are playing in the orchestra. But even I, with my limited hearing and even more limited musical knowledge, can tell when the brasses come in, and can often distinguish a French horn from a trombone even in the presence of a full orchestra. I doubt that I could distinguish a single note played on a harp from the same pitch played on a harpsichord, but the playing techniques are different enough to make identification possible when an entire melodic line is played. I can usually tell a clarinet from an oboe, and I think I can distinguish two of either playing in unison from one playing alone even when the whole orchestra joins them. Most orchestra musicians know which instrument played a single note even when a few play simultaneously. Jerry -- Engineering is the art of making what you want from things you can get. ¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯¯
Reply by ●May 6, 20082008-05-06
On May 6, 3:58�pm, "markt" <tak...@pericle.com> wrote:> >It takes a man to > >say he's wrong and move on. Then it's no big deal. Everyone is wrong > >sometimes. > > So, what exactly am I wrong about? �Oh, that you didn't say I claimed ICA > had solved the problem in the book... so what's your point about all this > ranting then? �Curious, since I never made any claims about it's validity. > > So, we're waiting, go ahead. > > While you're at it, look up the term "linear" with respect to mixtures. > I'm still laughing at that "point" you made. > > MarkMark, You make no sense. Are you fresh out of school or a summer hire? Off your medication? In love with ICA? Reread the posts and answer your own questions. I have never seen anyone in comp.dsp squirm as much as you. I have no clue what you are waiting for, but just stand there until it appears to you. Goodbye for good on this topic. Dirk
Reply by ●May 6, 20082008-05-06
>I have no clue what you are waiting for, but just stand there until it >appears to you.Your apology. Name one thing I said that is incorrect, btw...? And yes, you did imply that said the problem was solved in Hyvarinen's book.>Goodbye for good on this topic.Thank you. Btw all, enter "cocktail party problem" into wikipedia and you'll get redirected to the "Source Separation" page, which not surprisingly refers to both principal component analysis (PCA) as well as ICA as two methods for resolution to this issue. Source separation is a big deal, and means abound for approaching the issue. The wiki page also refers you to the Hyvarinen tutorial, which is also not a surprise since Nokia is sponsoring a lot of the research in Helsinki. The biggest problem with PCA is that it is based on correlations (variances), which offer no insight into independence. ICA allows one to exploit independence using higher-order statistics. Maximizing independence maximizes separation. If you have a vector space defined by a set of independent vectors, and your observations are a linear combination (scaling is allowed, btw) of these vectors, then there is one and only one solution that results in independent outputs. Not so with uncorrelated source vectors since it is very easy to define linear transformations of those vectors that are also uncorrelated, i.e., your solution may not be unique. Is ICA the end-all, be-all? No, and I said so in my second post. Is it worth investigating, particularly with this problem? Yes, and it even works, in spite of complaints to the contrary in this thread. Mark
Reply by ●May 7, 20082008-05-07
Sorry..I am a rookie here... I want to know what is the difference between 1. signals from two ppl from two different microphones(not present in the same scene) mixed linearly 2. two ppl's speech(spoken simultanously) recorded on same microphone .(does this resemble cocktail party!). I dont see any difference between these two problems. If i am wrong then what is the difference? regards Rajesh
Reply by ●May 7, 20082008-05-07
On 7 Mai, 07:04, rajesh <getrajes...@gmail.com> wrote:> Sorry..I am a rookie here... I want to know what is the difference > between > > � 1. signals from two ppl from two different microphones(not present > in the same scene) mixed linearly > > � 2. two ppl's speech(spoken simultanously) recorded on same > microphone .(does this resemble cocktail party!). > > �I dont see any difference between these two problems. If i am wrong > then what is the difference?If you look at the resulting single-channel data, there is litte difference. The one possible difference is that in your case 2 the two speakers' voices are affected by the 'same' [*] physical scenario (room impulse response) whereas in your case 1 there may be two different [*] room impulse responses involved. This may or may not have an impact on what DSP algorithms might work. Rune [*] Two people speaking simultaneously but at different positions in one room will cause different room impulse responses. However certain characteristics of those impulse responses (e.g. reverbrance time) will be similar. If two recordings are done in two different rooms with very different impulse responses, the different characteristics might (I'm speculating here!) just be enough to separate the two recordings.






